Senior HPC Infrastructure Engineer
firmus - Sydney, Australia
Posted Feb 9, 2026
Benefits
- Parental leave
- Not verified
- Non-birth-parent leave
- Not verified
- Family-building benefits
-
- Fertility benefits: Not verified
- Adoption assistance: Not verified
- Surrogacy assistance: Not verified
- Mental health support
- Not verified
- Relocation assistance
- Not verified
- Childcare support
- Not verified
- Learning budget
- Not verified
- Verification
- Not verified
- Salary
- Not verified
Was this benefit information wrong? Tell us.
Market context
- U.S. role benchmark (BLS OEWS)
- $116,543 U.S. median for this role
- Projected growth (BLS Employment Projections)
- +9.8% - Much faster than average
Matched to SOC 15-1252 - Software Engineering aggregate by role bucket.
Source: U.S. Bureau of Labor Statistics, OEWS, May 2024 and Employment Projections, 2024-2034.
Role
Schedule
- Shift type
- Not verified
- Weekend work
- Not verified
Application
- Cover letter
- Not verified
- Assessment
- Not verified
- Deadline
- Not stated
Where they hire
State eligibility is not yet verified.
About this role
Senior HPC Infrastructure Engineer Sydney, Australia Role Summary Firmus is seeking a highly skilled and driven Kubernetes HPC Engineer to join our Software Defined Infrastructure team. In this role, you will build high-performance, fault-tolerant, and reliable infrastructure to support bare-metal provisioning, performance benchmarking, and platform validation. You will be instrumental in ensuring the stability, performance, and continuous improvement of our complex and mission-critical bare-metal HPC GPU clusters. Key Responsibilities - Design and implement bare-metal provisioning workflows using Ironic and Kubernetes CRDs. - Deploy and manage GPU-enabled AI compute nodes with RDMA, InfiniBand, and RoCE networking. - Optimise Kubernetes and Slurm platforms for multi-node AI training performance, including NCCL, UCX, GPUDirect, and fabric tuning. - Implement Kubernetes primitives for GPU scheduling, isolation, and resource management models. - Design, deploy, and fine-tune Slurm GPU clusters with topology-aware configurations. - Develop and execute performance benchmarking workloads, including MLPerf, NCCL tests, microbenchmarks, and throughput/latency validation. - Establish observability across GPU, InfiniBand fabric, storage, and provisioning components. - Document architecture designs, operational procedures, and performance results. - Collaborate with L2 SRE engineers, site operations, and networking teams to ensure platform reliability, reproducibility, and performance. - Support hardware bring-up activities, including BIOS tuning, GPU topology verification, NUMA alignment, and PCIe/NVLink checks. - Contribute to continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks. - Contribute to the development of custom Kubernetes operators and intelligent orchestration frameworks that optimise AI workload performance for large-scale GPU cluster commissioning. Skills & Experience - Bachelor's or Master's
Read the full description at job-boards.greenhouse.io. FewerJobs shows a preview and links to the original posting.
Apply link not verified; last-live date unavailable.
What verified means
Verified means a displayed claim has field-level provenance to a source FewerJobs pulled: a government or employer source, or the original job posting. Posting-sourced facts are employer-stated and are labeled separately from government records.
Related jobs
-
Data Center Network Architect
Unisys CORP - Melbourne, VIC, Australia
-
Senior Solutions Engineer - Melbourne
Cisco - Melbourne, Australia
-
Senior Engineer - Mechanical / Structures
Northrop Grumman - Australia-Amberley
-
Strategy Consultant - Lab (AI and Frontier Tech)
Ivanhoe Electric INC - Melbourne