# Job array with pbspro

**URL:** <https://community.openpbs.org/t/job-array-with-pbspro/2079>\
**Category:** Users/Site Administrators\
**Created:** [April 16, 2020, 7:08am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079 "2020-04-16T07:08:35Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 16, 2020, 7:08am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/1 "2020-04-16T07:08:35Z")

</div>

Hi,

pbspro seems to treat a job array as a single job when attributes max\_queued and queued\_jobs\_threshold applied. are there any ways to treat each subjob of job array as one job so that max\_queued or queued\_jobs\_threshold can be applied to limit number of subjobs queuing? open source scheduler maui has this wonderful function.

thanks,

Sue

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [April 16, 2020, 8:39pm UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/2 "2020-04-16T20:39:41Z")

</div>

@sxy Could you please state the PBS Pro version you are running ?

---

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 16, 2020, 10:36pm UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/3 "2020-04-16T22:36:17Z")

</div>

V14.0.2

Thanks.

Sue

---

<div class="post-metadata">

**Author:** ![agrawalravi90](https://avatars.discourse-cdn.com/v4/letter/a/848f3c/32.png) [@agrawalravi90](https://community.openpbs.org/u/agrawalravi90)\
**Post date:** [April 16, 2020, 10:52pm UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/4 "2020-04-16T22:52:15Z")

</div>

The newer versions of PBS do treat subjobs as normal jobs for limits:

```
Qmgr: s s max_queued="[o:PBS_ALL=5]"
[ravi@pbspro ~]$ qsub -J 1-50 -- /bin/sleep 100
qsub: Maximum number of jobs already in complex
[ravi@pbspro ~]$ qsub -J 1-4 -- /bin/sleep 100
53[].pbspro
```

---

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 17, 2020, 12:04am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/5 "2020-04-17T00:04:20Z")

</div>

Qmgr: s s max\_queued="[o:PBS\_ALL=5]"  
[ravi@pbspro ~]$ qsub -J 1-50 – /bin/sleep 100  
qsub: Maximum number of jobs already in complex

can you run this command to show jobs’ status please,

[ravi@pbspro ~]$ qstat -1nt

thanks,

Sue

---

<div class="post-metadata">

**Author:** ![agrawalravi90](https://avatars.discourse-cdn.com/v4/letter/a/848f3c/32.png) [@agrawalravi90](https://community.openpbs.org/u/agrawalravi90)\
**Post date:** [April 17, 2020, 12:13am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/6 "2020-04-17T00:13:30Z")

</div>

> [@sxy](#):
>
> qstat -1nt

Output:

```
[ravi@pbspro ~]$ qsub -J 1-50 -- /bin/sleep 100
qsub: Maximum number of jobs already in complex
[ravi@pbspro ~]$ qstat -1nt
[ravi@pbspro ~]$ qstat -f
[ravi@pbspro ~]$ 
[ravi@pbspro ~]$ 
[ravi@pbspro ~]$

```

---

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 17, 2020, 1:06am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/7 "2020-04-17T01:06:27Z")

</div>

we have a routing queue as such

Queue defaultQ  
queue\_type = Route  
total\_jobs = 0  
state\_count = Transit:0 Queued:0 Held:0 Waiting:0 Running:0 Exiting:0 Begun:0  
route\_destinations = physics  
enabled = True  
started = True

execution queue, physics set as such,

Queue physics  
queue\_type = Execution  
Priority = 10  
total\_jobs = 0  
state\_count = Transit:0 Queued:0 Held:0 Waiting:0 Running:0 Exiting:0 Begun:0  
max\_queued = [u:PBS\_GENERIC=10]  
max\_queued = [u:sxy=10]  
max\_queued = [u:test=10]  
default\_chunk.Qlist = physics  
resources\_assigned.ncpus = 0  
resources\_assigned.nodect = 0  
enabled = True  
started = True  
queued\_jobs\_threshold = [u:PBS\_GENERIC=10]  
queued\_jobs\_threshold = [u:sxy=10]  
queued\_jobs\_threshold = [u:test=10]

we have 72 cores on queue, physics. with pbspro 14.0.1, if I run this

$qsub -J 1-200 – /bin/sleep 100  
$qsub -J 1-200 – /bin/sleep 100  
$qstat -1n

* * *

365[].headnode sxy physics STDIN – 1 1 – – B – –  
366[].headnode sxy defaultQ STDIN – 1 1 – – Q – –

for job 365[], 72 subjobs are running and others are queuing on queue, physics.  
I would expect that 10 subjobs were queuing in physics queue while others are queuing in defaultQ.  
also, job 366[] is waiting in queue, defaultQ until all 365 subjobs completes even there are free cores. then all subjobs entered into queue, physics at once.

could you test this scenario with version you run please?

nowdays, running job array is very popular on HPC systems. if one user submits a job array with large numbers of subjobs, perhaps thousands, it would cause"queue stuffing". we could use fairshare to limit jobs running for each user, but there would be overheads.

Thanks,

Sue

---

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 17, 2020, 1:58am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/8 "2020-04-17T01:58:28Z")

</div>

in your case, if you run

$ qsub – /bin/sleep 100

perhaps, you could also get this?

qsub: Maximum number of jobs already in complex

Sue

---

<div class="post-metadata">

**Author:** ![agrawalravi90](https://avatars.discourse-cdn.com/v4/letter/a/848f3c/32.png) [@agrawalravi90](https://community.openpbs.org/u/agrawalravi90)\
**Post date:** [April 17, 2020, 2:07am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/9 "2020-04-17T02:07:32Z")

</div>

Ok, i tried to create a similar setup:

```
[ravi@pbspro ~]$ qstat -Qf
Queue: physics
    queue_type = Execution
    Priority = 10
    total_jobs = 0
    state_count = Transit:0 Queued:0 Held:0 Waiting:0 Running:0 Exiting:0 Begun
	:0 
    max_queued = [u:PBS_GENERIC=10]
    max_queued = [u:ravi=10]
    enabled = True
    started = True

Queue: defaultQ
    queue_type = Route
    total_jobs = 2
    state_count = Transit:0 Queued:2 Held:0 Waiting:0 Running:0 Exiting:0 Begun
	:0 
    route_destinations = physics
    enabled = True
    started = True

Queue: defaultQ
    queue_type = Route
    total_jobs = 0
    state_count = Transit:0 Queued:0 Held:0 Waiting:0 Running:0 Exiting:0 Begun
	:0 
    route_destinations = physics
    enabled = True
    started = True

Submitted 2 jobs as user 'ravi':
[ravi@pbspro ~]$ qsub -J 1-200 -q defaultQ -- /bin/sleep 1000
60[].pbspro
[ravi@pbspro ~]$ qsub -J 1-200 -q defaultQ -- /bin/sleep 1000
61[].pbspro

```

They were both queued in defaultQ:  
[ravi@pbspro ~]$ qstat -1n

```
pbspro: 
                                                            Req'd Req'd Elap
Job ID Username Queue Jobname SessID NDS TSK Memory Time S Time
--------------- -------- -------- ---------- ------ --- --- ------ ----- - -----
60[].pbspro ravi defaultQ STDIN -- 1 1 -- -- Q -- -- 
61[].pbspro ravi defaultQ STDIN -- 1 1 -- -- Q -- -- 

```

So, I guess PBS doesn’t allow an array job to be queued into the execution queue unless it can completely fit it, so they both stay in the routing queue. If I submit a smaller job from the same user, that does routed to physics and starts running:

```
[ravi@pbspro ~]$ qsub -J 1-5 -q defaultQ -- /bin/sleep 1000
62[].pbspro
[ravi@pbspro ~]$ qstat -1n

pbspro: 
                                                            Req'd Req'd Elap
Job ID Username Queue Jobname SessID NDS TSK Memory Time S Time
--------------- -------- -------- ---------- ------ --- --- ------ ----- - -----
60[].pbspro ravi defaultQ STDIN -- 1 1 -- -- Q -- -- 
61[].pbspro ravi defaultQ STDIN -- 1 1 -- -- Q -- -- 
62[].pbspro ravi physics STDIN -- 1 1 -- -- B -- --
```

---

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 17, 2020, 11:48am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/10 "2020-04-17T11:48:10Z")

</div>

[quote=“agrawalravi90, post:4, topic:2079”

job 60[] and 61[] above will never run? I would expect that 5 subjobs routed to queue physics would run. we strongly suggest the future pbspro would include such function for subjobs in job array as more and more users are running job array on HPC.  
Apparently this function has many advantages such as speeding up the scheduling cycle in execution queues and preventing single user of stuffing execution queues.  
open source scheduler maui has this function for many years and I am surprised that pbspro has not adopted it yet.

thanks,

Sue

---

<div class="post-metadata">

**Author:** ![agrawalravi90](https://avatars.discourse-cdn.com/v4/letter/a/848f3c/32.png) [@agrawalravi90](https://community.openpbs.org/u/agrawalravi90)\
**Post date:** [April 17, 2020, 2:44pm UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/11 "2020-04-17T14:44:13Z")

</div>

> [@sxy](#):
>
> job 60 and 61 above will never run?

Well, they cross the max\_queued limit, so ya, they won’t ever get queued to run. Would it be unreasonable to ask your users to submit smaller job arrays? As long as users submit arrayjobs within the limits, your use case will be achieved.

But thanks for mentioning this, we can analyze further whether it makes sense to enhance job arrays in PBS to be able to queue some of it into the execution queue if max\_queued is set.

---

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 18, 2020, 1:43am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/12 "2020-04-18T01:43:07Z")

</div>

fundamental idea is that an array job as a subjob in job array should be treated equally as a normal job so that any functions with pbspro apply to normal jobs should also apply to array jobs. could you please pass my comments under this topic to your manager?

thanks,

Sue

---

<div class="post-metadata">

**Author:** ![agrawalravi90](https://avatars.discourse-cdn.com/v4/letter/a/848f3c/32.png) [@agrawalravi90](https://community.openpbs.org/u/agrawalravi90)\
**Post date:** [April 19, 2020, 4:02am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/13 "2020-04-19T04:02:54Z")

</div>

I think the new version of PBS does treat subjobs pretty much the same way as normal jobs. Can you please be more specific about what differences you see between a subjob and a normal job in PBS? The limits thing, as I explained earlier, has been fixed in the latest versions of PBS. Is there any other difference that you are concerned about?

---

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 19, 2020, 4:20am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/14 "2020-04-19T04:20:12Z")

</div>

Did you use the latest version in your example in our conversions?

Sue

---

<div class="post-metadata">

**Author:** ![agrawalravi90](https://avatars.discourse-cdn.com/v4/letter/a/848f3c/32.png) [@agrawalravi90](https://community.openpbs.org/u/agrawalravi90)\
**Post date:** [April 19, 2020, 3:45pm UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/15 "2020-04-19T15:45:25Z")

</div>

> [@sxy](#):
>
> Did you use the latest version in your example in our conversions?

I used the master branch in the examples above.

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [April 20, 2020, 11:52am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/16 "2020-04-20T11:52:04Z")

</div>

@sxy

If you are using 14.X then you might face the issue and it might not work the way you like.  
If you download 19.x or the latest from the master as @agrawalravi90 suggested, that would work the way you like .

The behaviour you have seen in 14.x is a bug and it is fixed in version 19.x or the master branch.

---

<div class="post-metadata">

**Author:** ![sxy](https://avatars.discourse-cdn.com/v4/letter/s/c89c15/32.png) [@sxy](https://community.openpbs.org/u/sxy)\
**Post date:** [April 22, 2020, 1:40am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/18 "2020-04-22T01:40:39Z")

</div>

Hi,

The newer versions of PBS do treat subjobs as normal jobs for limits:

```auto
Qmgr: s s max_queued="[o:PBS_ALL=5]"
[ravi@pbspro ~]$ qsub -J 1-50 -- /bin/sleep 100
qsub: Maximum number of jobs already in complex

```

in your case, no jobs ran. but I would expect first 5 jobs should start to run if subjobs are treated in the same way as normal jobs.

thanks,

Sue

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [April 22, 2020, 7:07am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/19 "2020-04-22T07:07:21Z")

</div>

Hi,  
Please check the admin guide : [https://www.altair.com/pdfs/pbsworks/PBS19.2.3\_BigBook.pdf](https://www.altair.com/pdfs/pbsworks/PBS19.2.3_BigBook.pdf) on section  
Table 4-1: Server Attributes Involved in Scheduling  
**max\_queued:** The **maximum number of jobs allowed to be queued or running** in the partition(s) managed by a scheduler. Can be specified for users, groups, or all.

Thank you

---

<div class="post-metadata">

**Author:** ![bhroam](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.openpbs.org/bhroam/32/20_2.png) [@bhroam](https://community.openpbs.org/u/bhroam)\
**Post date:** [April 22, 2020, 6:16pm UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/20 "2020-04-22T18:16:57Z")

</div>

A job array isn’t a collection of jobs. It’s one “job” that can spawn many jobs from it. This means that a subjob doesn’t exist until it starts running. We can’t detach N subjobs from the parent to move into the exec queue. They don’t exist yet. The whole job array has to move.

With this limitation, fixing it as you suggest is quite difficult.

I am not fully up to date on job arrays though. This might be easier than I think. Hopefully our job array expert (@Shrini-h) can weigh in.

Bhroam

---

<div class="post-metadata">

**Author:** ![Shrini-h](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.openpbs.org/shrini-h/32/70_2.png) [@Shrini-h](https://community.openpbs.org/u/Shrini-h)\
**Post date:** [April 30, 2020, 4:48am UTC](https://community.openpbs.org/t/job-array-with-pbspro/2079/21 "2020-04-30T04:48:15Z")

</div>

Sorry for the delay in replying

I think @sxy has a valid expectation, unfortunately the current architecture doesn’t allow it.

I also agree with @Bhroam that its not an easy enhancement, but will be an interesting one to solve. there could be many ways to architect 1. decoupling parent job from queue (make it sort of global) and let subjobs freely move around queues, 2. Routing job array can spawn smaller dependent job arrays at the execution queue… etc

Thank You

[Next page](https://community.openpbs.org/t/job-array-with-pbspro/2079.md?page=2)
