Slurm HPC Cluster
Designed, planned and implemented the first and only Slurm cluster at HEIA-FR / iCoSys to schedule NVIDIA GPUs (GH200, A40, A6000) used by many data scientists. I configured the full CUDA environment and the NVIDIA Container Toolkit, and mixed bare-metal GPUs with virtual GPUs (vGPU) running in VMs on a vSphere cluster, with cloud bursting from those VMs. The stack is deployed automatically, precisely monitored, and gives users access to three separate storage spaces.