This shows you the differences between two versions of the page.
| Both sides previous revision Previous revision Next revision | Previous revision | ||
|
quick:walltime [2021/03/26 15:31] agill [How to Estimate Walltime] |
quick:walltime [2021/12/09 16:42] (current) |
||
|---|---|---|---|
| Line 3: | Line 3: | ||
| As most CHPC users will know, there is a limit on how long a job submitted to any queue on Lengau will be allowed to run for, as is specified in the user policy document. | As most CHPC users will know, there is a limit on how long a job submitted to any queue on Lengau will be allowed to run for, as is specified in the user policy document. | ||
| For instance, at present, jobs submitted to the normal queue cannot run for more than 48 hours. | For instance, at present, jobs submitted to the normal queue cannot run for more than 48 hours. | ||
| + | In your script, you'd specify it something like this: | ||
| - | It might be tempting, therefore, to specify this for the maximum allowed walltime for the queue, and not worry about how long the job will actually run for. However, this is not actually a good idea. | + | |
| + | |||
| + | It might be tempting, therefore, to specify this for the maximum allowed walltime for the queue, and not worry about how long the job will actually run for. However, this is not actually a good idea and may waste your time by delaying the start of your jobs. | ||
| When you submit a job using qsub, it gets passed to a piece of cluster-management software called the job-scheduler (we use PBS-Pro on lengau at present). This software looks at the requested walltime, cores, nodes and other resources requested for each job, and attempts to fit as many jobs into the available resources on the cluster as soon as is reasonable, so that no one job has too long a wait. In effect, you might say that it is a bit like playing tetris with the jobs. | When you submit a job using qsub, it gets passed to a piece of cluster-management software called the job-scheduler (we use PBS-Pro on lengau at present). This software looks at the requested walltime, cores, nodes and other resources requested for each job, and attempts to fit as many jobs into the available resources on the cluster as soon as is reasonable, so that no one job has too long a wait. In effect, you might say that it is a bit like playing tetris with the jobs. | ||
| Line 13: | Line 16: | ||
| ===== How to Estimate Walltime ===== | ===== How to Estimate Walltime ===== | ||
| So how do you make a better estimate of your walltime? | So how do you make a better estimate of your walltime? | ||
| - | Well, when you begin doing a new type of simulation, you should do a few tests to see how long a few representative jobs will run for. In this case, it is all right if your walltime estimate is inaccurate - because you won't have to do this often. | + | Well, when you begin doing a new type of simulation, you should do a few tests to see how long a few representative jobs will run for. In this case, it is all right if your walltime estimate is inaccurate - because you won't have to do this often. |
| ===== Using Scaling to Decrease Walltime ===== | ===== Using Scaling to Decrease Walltime ===== | ||
| - | Something else which can help you use shorter walltimes is to test the scaling of your code. This is dealt with elsewhere on the wiki ,but briefly speaking, are you using as many cores as you can efficiently use to speed up your simulation? And are you perhaps using too many cores, which can actually slow things down? Again, when you begin running a new piece of software or a new type of calculation on the cluster, do a few runs to experiment on how much you will benefit from using more cores, and at what point you stop benefiting by requesting more cores. Going back to our Tetris analogy, if you can turn a long shape sideways and can lay it flat, you are usually better off than trying to stand it on its end. Similarly, if you can turn your long-duration, | + | Something else which can help you use shorter walltimes is to test the scaling of your code. This is dealt with elsewhere on the wiki ([[scaling: |
| ===== Long Walltime Jobs and Checkpointing ===== | ===== Long Walltime Jobs and Checkpointing ===== | ||