Domain Architect - AI Compute - 12M CONTRACT - PERM
Job description
Salary: Melbourne, 12M Contract. Attractive Daily Rates
ACT NOW: Opportunity to join a market leading System Integrator delivering Multi-Billion-Dollar AI Data Centre Upgrade Programs. Work with the New and Emerging technology such as NVIDIA in an Org organisation at the forefront of delivering AI Data Centre Upgrade Programs, Delivering GPUs at scale.We have a (12-Month Contract - Permanent) Domain Architect – AI Compute roles available.
You will join a specialist team delivering high-performance GPU and AI infrastructure for enterprise clients.
This is a hands-on architecture and delivery role requiring strong NVIDIA GPU, HPC and Linux expertise.
You will own the Compute layer across complex AI infrastructure deployments, working across architecture, cluster provisioning, automation, performance optimisation and Day 2 operations.
Key
Responsibilities
- Design and deliver NVIDIA SuperPOD, BasePOD, NVL72, HGX, MGX and Cisco AI Factory environments.
- Lead bare-metal GPU cluster provisioning using NVIDIA Base Command Manager (BCM) and Zero Touch Provisioning.
- Configure and optimise Slurm, NVIDIA Run, Kueue and/or Volcano.
- Implement GPU monitoring using NVIDIA DCGM and validate performance using NCCL, HPL and HPCG.
- Optimise Linux environments for high-performance AI/HPC workloads, including kernel, driver and system tuning.
- Develop infrastructure automation using Python and Ansible.
- Integrate compute environments with InfiniBand/RoCEv2, storage, identity and enterprise platforms.
- Provide technical leadership to clients and support solution design, scoping and pre-sales activities.
- Strong HPC Architecture & Engineering Experience
- Strong Experience in Data Centre - Compute
- Strong hands-on experience with NVIDIA GPU/AI infrastructure and HPC environments.
- Deep knowledge of NVIDIA Hopper, Grace Hopper, Blackwell and Grace Blackwell platforms.
- Experience with NVL72, NVSwitch, CUDA, cuDNN and NCCL.
- Advanced Linux systems engineering skills across Ubuntu and/or RHEL.
- Strong Python and Ansible automation experience.
- Understanding of GPU/PCIe/NUMA topology and high-performance compute architectures.
- NVIDIA BCM/Mission Control, Slurm, Run, Kubernetes/OpenShift, InfiniBand NDR/HDR, RoCEv2, Rafay, KubeFlow, OpenShift AI, liquid cooling and AWS/Azure/GCP/OCI.
- Experience within a System Integrator, MSP, hyperscaler or specialist AI infrastructure environment is highly regarded.
This is a great opportunity - To apply, please submit your CV via the portal by clicking the APPLY NOW button below.
You can also contact Charlie directly at: [email protected]
Charlie Molino
0450 253 077
Northbridge IT Recruitment
For this and other opportunities please visit:
[external link]
Skills mentioned
Apply for this job
Use the application link supplied with this listing to apply to Northbridge Recruitment. Check the destination before entering personal information.
