Mirantis

HPC Network Engineer

Tokyo, Tokyo, JapanRemoteFull timePosted 10 days ago
Apply on Mirantis →

Sign into see who you know at Mirantis.

About MirantisMirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. We are looking for an HPC Network Engineer to join our global team responsible for managing and supporting high-performance networking environments for large-scale AI infrastructure. You will help ensure the performance, reliability, and security of critical network fabrics, working closely with distributed compute and storage teams to support modern AI workloads.This is an opportunity to work hands-on with some of the most advanced InfiniBand and Ethernet networking infrastructure in production today, while developing deep expertise in AI infrastructure and high-performance computing (HPC) technologies.Responsibilities:Support the design, deployment, configuration, and maintenance of InfiniBand and Ethernet network infrastructures.Troubleshoot complex network issues, including connectivity, latency, routing, and performance degradation across hybrid environments.Manage and optimize high-performance network components, such as switches, Host Channel Adapters (HCAs), subnet managers, and fabric configurations.Implement, manage, and troubleshoot network security and firewall technologies (e.g., Fortinet solutions like FortiGate, VPNs).Monitor network health, perform performance tuning, and assist in capacity planning for HPC and AI networking systems.Collaborate with compute, storage, and platform teams to seamlessly support HPC and AI workloads.Participate in incident response, on-call support activities, and drive long-term operational improvements.Document network architectures, configurations, and operational procedures, and adopt automation tools (e.g., Ansible, Terraform) to streamline workflows. Proven experience in networking, system engineering, or data center infrastructure roles.Solid understanding of networking fundamentals, including TCP/IP, routing protocols (BGP, OSPF), switching, VLANs, QoS, and network design.Hands-on experience or strong familiarity with Linux operating systems and command-line tools.Effective verbal and written communication skills in English, with the ability to collaborate in a global team environment.Strong analytical and problem-solving skills to diagnose a...

Also hiring in