Team Lead, Platform Engineering
Overview: We are building a modern internal platform that enables application, data, and AI teams to build, deploy, and scale software efficiently and securely. We are looking for an experienced Principal Platform Engineer to drive the development of our Internal Developer Platform (IDP). The ideal candidate started their career as a software engineer, later transitioned into platform engineering, and has grown into a technical leader capable of guiding teams and building scalable platforms. This role combines hands-on engineering with leadership, ensuring the platform delivers a great developer experience while maintaining reliability, governance, and scalability. You will provide guidance and mentorship to a team of platform engineers while working closely with application teams, data teams, cloud operations, and security teams to enable modern cloud-native, Data and AI workloads. This is a hybrid position based in our Toronto office. What you will do: Platform strategy and engineering Lead the design and evolution of a scalable Internal Developer Platform (IDP) that enables self-service infrastructure and developer productivity. Architect and operate production-grade Kubernetes platforms supporting microservices, data workloads, and AI/ML systems. Implement GitOps-based deployment workflows using Argo CD. Build and maintain developer portals using Backstage to streamline service onboarding and platform access. Implement event-driven autoscaling using KEDA for scalable workloads. Enforce platform governance through policy-as-code using Kyverno. Build reusable platform components, templates, and golden paths that simplify development workflows. Manage and maintain all platform tools CI/CD and platform automation Design and standardize CI/CD pipelines for both modern cloud-native applications and legacy development environments. Implement GitOps and infrastructure-as-code practices to ensure consistent, automated platform operations. Enable engineering teams with self-service deployment and infrastructure capabilities. Observability and reliability Establish end-to-end observability practices across the platform using Datadog. Implement monitoring, logging, tracing, and reliability practices to ensure highly resilient platform services. AI platform enablement Build platform capabilities that support AI/ML experimentation, model training, and inference workloads. Collaborate with data scientists and ML engineers to enable scalable and reliable AI infrastructure. Platform product management Operate the platform with a platform-as-a-product mindset, continuously improving capabilities based on developer needs and feedback. Define and maintain platform roadmaps, standards, and service offerings that align with engineering and business goals. Measure and improve developer productivity and plat...