# GPU mapping in PBS

**URL:** <https://community.openpbs.org/t/gpu-mapping-in-pbs/3591>\
**Category:** Users/Site Administrators\
**Created:** [June 27, 2023, 8:33pm UTC](https://community.openpbs.org/t/gpu-mapping-in-pbs/3591 "2023-06-27T20:33:12Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![berceanu](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.openpbs.org/berceanu/32/651_2.png) [@berceanu](https://community.openpbs.org/u/berceanu)\
**Post date:** [June 27, 2023, 8:33pm UTC](https://community.openpbs.org/t/gpu-mapping-in-pbs/3591/1 "2023-06-27T20:33:12Z")

</div>

I am running an MPI + CUDA HPC code on a system using `Open PBS` with multiple nodes, each node having 8 NVIDIA GPUs.

On a `SLURM` cluster, I can use “pinning” in order to assign the MPI  
rank to the CPUs that are closest to each of the GPUs.

For simplicity, let’s assume we have a node with 4 GPUs and 16 CPUs (or cores), and we want to pin 4 MPI tasks such that each task is associated with one GPU and 4 cores that are closest to it. Here’s a simplified version of how I might go about doing it:

```auto
#SBATCH --nodes=1
#SBATCH --ntasks=4
#SBATCH --cpus-per-task=4
#SBATCH --gres=gpu:4
#SBATCH --cpu-bind=cores
#SBATCH --gpu-bind=map_gpu:0,1,2,3
mpirun ./my_application

```

An alternative approach (which however does not minimize CPU-GPU latency) is to use `CUDA_VISIBLE_DEVICES`, like so:

```auto
#!/bin/bash
#SBATCH --ntasks=4
#SBATCH --gres=gpu:4

mpirun -np 4 -x CUDA_VISIBLE_DEVICES=$SLURM_LOCALID ./my_application

```

How can I do this using `PBS` for maximising the performance of the code?

---

<div class="post-metadata">

**Author:** ![alexis.cousein](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.openpbs.org/alexis.cousein/32/265_2.png) [@alexis.cousein](https://community.openpbs.org/u/alexis.cousein)\
**Post date:** [July 3, 2023, 11:11pm UTC](https://community.openpbs.org/t/gpu-mapping-in-pbs/3591/2 "2023-07-03T23:11:28Z")

</div>

Use the cgroup hook and enable vnode\_per\_numa\_node, it will make the scheduler aware of the topology. But if you’re spanning more than one socket then you still have to discover what process to pin where and which GPU to use from that process.

---

<div class="post-metadata">

**Author:** ![berceanu](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.openpbs.org/berceanu/32/651_2.png) [@berceanu](https://community.openpbs.org/u/berceanu)\
**Post date:** [July 4, 2023, 9:17am UTC](https://community.openpbs.org/t/gpu-mapping-in-pbs/3591/3 "2023-07-04T09:17:49Z")

</div>

Thank you for the information!  
There are indeed 2 sockets per node, 4 GPUs per socket. Do you think you could maybe provide some sample code to make this clearer?
