User Tools

Site Tools


guide:gpu

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revision Previous revision
Next revision
Previous revision
guide:gpu [2020/09/14 11:32]
kevin [GPU Queues]
guide:gpu [2025/08/07 12:51] (current)
kevin
Line 11: Line 11:
 | ''gpu2003''  |  36 | 3× Nvidia V100 16GB  | PCIe  | | ''gpu2003''  |  36 | 3× Nvidia V100 16GB  | PCIe  |
 | ''gpu2004''  |  36 | 3× Nvidia V100 16GB  | PCIe  | | ''gpu2004''  |  36 | 3× Nvidia V100 16GB  | PCIe  |
-| ''gpu2005''  |  36 | 3× Nvidia V100 32GB  | PCIe  | +| ''gpu2005''  |  36 | 3× Nvidia V100 **32GB**  | PCIe  | 
-| ''gpu2006''  |  36 | 3× Nvidia V100 32GB  | PCIe  | +| ''gpu2006''  |  36 | 3× Nvidia V100 **32GB**  | PCIe  | 
-| ''gpu4001''  |  40 | 4× Nvidia V100 16GB  | NVlink +| ''gpu4001''  |  40 | 4× Nvidia V100 16GB  | **NVlink**  | 
-| ''gpu4002''  |  40 | 4× Nvidia V100 16GB  | NVlink +| ''gpu4002''  |  40 | 4× Nvidia V100 16GB  | **NVlink**  | 
-| ''gpu4003''  |  40 | 4× Nvidia V100 16GB  | NVlink  |+| ''gpu4003''  |  40 | 4× Nvidia V100 16GB  | **NVlink**  |
  
 > Jobs that require 1, 2 or 3 GPUs can be allocated to any node, and will share the node if the job does not use all the GPU devices on that node.  Jobs that require 4 GPUs can only be allocated to ''gpu4*'' nodes and will not be shared, obviously. > Jobs that require 1, 2 or 3 GPUs can be allocated to any node, and will share the node if the job does not use all the GPU devices on that node.  Jobs that require 4 GPUs can only be allocated to ''gpu4*'' nodes and will not be shared, obviously.
Line 27: Line 27:
 ====Access==== ====Access====
  
-Access to these GPU node is by PI application only though the CHPC [[https://users.chpc.ac.za/helpdesk/|Helpdesk]].+Principle Investigators apply for GPU access for their Research Programme members through the CHPC [[https://users.chpc.ac.za/helpdesk/|Helpdesk]].  RP members may not apply directly. E-mailed applications will not be considered
  
 ====Allocation===== ====Allocation=====
Line 52: Line 52:
  
 ^  Queue name  ^  Max. CPUs  ^  Max. GPUs  ^  PBSPro options  ^ Comments  ^ ^  Queue name  ^  Max. CPUs  ^  Max. GPUs  ^  PBSPro options  ^ Comments  ^
-|  **gpu_1**  |  10 |  1 | ''-q gpu_1''\\ ''-l ncpus=9:ngpus=1''  | Access one GPU device only per job.  | +|  **gpu_1**  |  |  1 | ''-q gpu_1''\\ ''-l select=1:ncpus=9:ngpus=1''  | Access one GPU device only per job.  | 
-|  **gpu_2**  |  20 |  2 | ''-q gpu_2''\\ ''-l ncpus=18:ngpus=2''  | Access two GPU devices per job.  | +|  **gpu_2**  |  18 |  2 | ''-q gpu_2''\\ ''-l select=1:ncpus=18:ngpus=2''  | Access two GPU devices per job.  | 
-|  **gpu_3**  |  36 |  3 | ''-q gpu_3''\\ ''-l ncpus=30:ngpus=3''  | Access three GPU devices per job.  | +|  **gpu_3**  |  36 |  3 | ''-q gpu_3''\\ ''-l select=1:ncpus=30:ngpus=3''  | Access three GPU devices per job.  | 
-|  **gpu_4**  |  40 |  4 | ''-q gpu_4''\\ ''-l ncpus=40:ngpus=4''  | Access four GPU devices on NVLink nodes.  |+|  **gpu_4**  |  40 |  4 | ''-q gpu_4''\\ ''-l select=1:ncpus=40:ngpus=4''  | Access four GPU devices on NVLink nodes.  |
  
-Note the ''ncpus'' parameters that should be set to match the number of GPU devices you need.+Note the ''ncpus'' parameters above is the maximum that should be set to match the number of GPU devices you need.
  
 ====GPU Queue Limits==== ====GPU Queue Limits====
Line 73: Line 73:
  
 <code bash> <code bash>
-qsub -I -q gpu_1 -P PRJT1234+qsub -I -q gpu_1 -P PRJT1234 -l select=1:ncpus=9:ngpus=1
 </code> </code>
 **NB:** Replace //''PRJT1234''// with **//your//** project number. **NB:** Replace //''PRJT1234''// with **//your//** project number.
Line 85: Line 85:
 #PBS -N nameyourjob #PBS -N nameyourjob
 #PBS -q gpu_1 #PBS -q gpu_1
-#PBS -l ncpus=10:ngpus=1+#PBS -l select=1:ncpus=4:ngpus=1
 #PBS -P PRJT1234 #PBS -P PRJT1234
 #PBS -l walltime=4:00:00 #PBS -l walltime=4:00:00
-#PBS -o /mnt/lustre/users/USERNAME/cuda_test/test1.out 
-#PBS -e /mnt/lustre/users/USERNAME/cuda_test/test1.err 
 #PBS -m abe #PBS -m abe
 #PBS -M your.email@address #PBS -M your.email@address
Line 97: Line 95:
 echo echo
 echo `date`: executing CUDA job on host ${HOSTNAME} echo `date`: executing CUDA job on host ${HOSTNAME}
 +echo
 +echo Available GPU devices: $CUDA_VISIBLE_DEVICES
 echo echo
  
Line 102: Line 102:
 ./hello_cuda ./hello_cuda
 </code> </code>
 +
 +As usual, replace ''PRJT1234'' with your group's project name, ''your.email@address'' with your email address, and ''USERNAME'' with your cluster user name.
 +
 +====GPU Memory====
 +
 +Most of the V100 GPUs only have 16GiB of memory.  Two nodes each have three V100s with 32GiB of RAM.  To access these it is necessary to specify the exact node:
 +
 +  #PBS -l select=1:host=gpu2005:ncpus=9:ngpus=1
 +
 +or
 +
 +  #PBS -l select=1:host=gpu2006:ncpus=9:ngpus=1
 +
 +> Note that asking for a specific node is likely to lead to longer queue times as your job has to wait until that node becomes available.
 +  
  
 =====Compiling GPU Code===== =====Compiling GPU Code=====
Line 107: Line 122:
 The Nvidia V100 GPUs are programmed using the **CUDA** development tools. The Nvidia V100 GPUs are programmed using the **CUDA** development tools.
  
-To build a CUDA code (library or application) for the GPU nodes requires loading the appropriate CUDA module before compiling.  The CUDA runtime tools are already installed on all GPU nodes and won't need to be loaded specifically unless you require (for some reason) a different versionThe V100 GPUs have Volta architecture coresCUDA applications built using CUDA Toolkit versions 2.1 through 8.0 are compatible with Volta as long as they are built to include PTX versions of their kernelsTo test that PTX JIT is working for your application, you can do the following: +To build a CUDA code (library or application) for the GPU nodes requires loading the appropriate CUDA module before compiling.    The current CUDA modules are
-Download and install the latest driver from [[http://www.nvidia.com/drivers]]+ 
-Set the environment variable CUDA_FORCE_PTX_JIT=1+<code> 
-Launch your application+chpc/cuda/11.2/PCIe/11.2 
-When starting a CUDA application for the first time with the above environment flag, the CUDA driver will JIT-compile the PTX for each CUDA kernel that is used into native cubin code.+chpc/cuda/11.2/SXM2/11.2 
 +chpc/cuda/11.5.1/PCIe/11.5.1 
 +chpc/cuda/11.6/PCIe/11.6 
 +chpc/cuda/11.6/SXM2/11.6 
 +chpc/cuda/12.0/12.0 
 +</code> 
 + 
 +with version 12.0 the most recent version. 
 + 
 +Note that the 11.x version modules are available in two types:   
 +    * the ''PCIe'' version is for PCIe bus nodes: ''gpu200x'' 
 +    * the ''SXM2'' version is for the SXM2 bus nodes: ''gpu400x''
  
-If you set the environment variable above and then launch your program and it works properly, then you have successfully verified Volta compatibility. 
  
-Note: Be sure to unset the CUDA_FORCE_PTX_JIT environment variable when you are done testing. 
  
 ====Further Reading==== ====Further Reading====
/app/dokuwiki/data/attic/guide/gpu.1600075957.txt.gz · Last modified: 2021/12/09 16:42 (external edit)