Lead Site Reliability Engineer
ABOUT USWe’re the world’s leading provider of secure financial messaging services, headquartered in Belgium. We are the way the world moves value – across borders, through cities and overseas. No other organisation can address the scale, precision, pace and trust that this demands, and we’re proud to support the global economy. We’re unique too. We were established to find a better way for the global financial community to move value – a reliable, safe and secure approach that the community can trust, completely. We’re always striving to be better and are constantly evolving in an ever-changing landscape, without undermining that trust. Five decades on, our vibrant community reflects the complexity and diversity of the financial ecosystem. We innovate diligently, test exhaustively, then implement fast. In a connected and exciting era, our mission has never been more relevant. Swift now has a presence in 200+ countries and legal territories to serve a community of more than 12,000 banks and financial institutions. The ideal candidate for this position will possess a strong technical foundation in software development in enterprise environment. You will be responsible for design, develop, test, deliver and support for large-scale of infrastructure such as stability, patching and upgrade. Within this infrastructure is running large-scale of data pipelines with big data technologies such as Elasticsearch, Logstash, Kibana, Kafka etc. He/She is able to build strong partnerships with internal customers and other delivery organizations inside SWIFT.Job DescriptionWhat to expect:Work through all phases of the system administration life cycle, including capacity planning, architecture design, compliance, deployment & configuration, monitoring, and incident management.Develop automation scripts, infrastructure as code, and tooling using industry best practices to improve system reliability, reduce manual effort, and enable self-service.Review system architectures design, deployment strategies, observability setups, and operational documentation to ensure reliability and operational excellence.Analyze production issues, identify root causes, and implement long-term reliability improvements through automation, monitoring, and architectural enhancements.Work collaboratively with other team members and provide guidance to more junior team members.Organize an efficient handover through high quality documentation and training.Automate the deployment and operation of multi-tenant infrastructure, handling tasks that ensure system resilience and availability.Develop and maintain monitoring tools, dashboards, and self-healing mechanisms.Participate in on-call rotations, conduct blameless postmortems, and drive continuous learning.Work closely with developers, product teams, and engineering stakeholders to troubleshoot issues, improve systems, and integrate reliability improvementsCapable of providing accurate project estimates and strate...