# Job stuck in queue, multiple servers

**URL:** <https://community.openpbs.org/t/job-stuck-in-queue-multiple-servers/3218>\
**Category:** Users/Site Administrators\
**Created:** [July 20, 2022, 12:12pm UTC](https://community.openpbs.org/t/job-stuck-in-queue-multiple-servers/3218 "2022-07-20T12:12:08Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Blackest](https://avatars.discourse-cdn.com/v4/letter/b/5e9695/32.png) [@Blackest](https://community.openpbs.org/u/Blackest)\
**Post date:** [July 20, 2022, 12:12pm UTC](https://community.openpbs.org/t/job-stuck-in-queue-multiple-servers/3218/1 "2022-07-20T12:12:08Z")

</div>

Greetings,

I have installed openpbs on a workstation (a single machine run everything, it is possible to remotely access through ssh).  
I managed to send jobs just fine until a couple of months ago. I haven’t used it for a while and now that I’m back the job I sent are stuck in queue. There are no other jobs running. If I delete the jobs no error files are created. I checked logs and everything seems fine.

I noted this detail though, when I input pbsnodes -av I have two identical servers with different status (stale and free)

sysadmin@Precision-7920-Tower:~/testVASP/newtest$ pbsnodes -av  
precision-7920-tower  
Mom = precision-7920-tower  
ntype = PBS  
state = Stale  
pcpus = 20  
resources\_available.arch = linux  
resources\_available.host = precision-7920-tower  
resources\_available.mem = 97495476kb  
resources\_available.ncpus = 20  
resources\_available.vnode = precision-7920-tower  
resources\_assigned.accelerator\_memory = 0kb  
resources\_assigned.hbmem = 0kb  
resources\_assigned.mem = 0kb  
resources\_assigned.naccelerators = 0  
resources\_assigned.ncpus = 0  
resources\_assigned.vmem = 0kb  
resv\_enable = True  
sharing = default\_shared  
license = l  
last\_state\_change\_time = Wed Jul 20 14:45:18 2022  
last\_used\_time = Tue Jan 18 22:15:56 2022

Precision-7920-Tower  
Mom = precision-7920-tower  
ntype = PBS  
state = free  
pcpus = 20  
resources\_available.arch = linux  
resources\_available.host = precision-7920-tower  
resources\_available.mem = 97495476kb  
resources\_available.ncpus = 20  
resources\_available.vnode = Precision-7920-Tower  
resources\_assigned.accelerator\_memory = 0kb  
resources\_assigned.hbmem = 0kb  
resources\_assigned.mem = 0kb  
resources\_assigned.naccelerators = 0  
resources\_assigned.ncpus = 0  
resources\_assigned.vmem = 0kb  
resv\_enable = True  
sharing = default\_shared  
license = l  
last\_state\_change\_time = Wed Jul 20 14:45:18 2022  
last\_used\_time = Wed Jul 20 13:16:42 2022

I wonder if this could be the issue.

Thanks!

---

<div class="post-metadata">

**Author:** ![Blackest](https://avatars.discourse-cdn.com/v4/letter/b/5e9695/32.png) [@Blackest](https://community.openpbs.org/u/Blackest)\
**Post date:** [July 20, 2022, 12:42pm UTC](https://community.openpbs.org/t/job-stuck-in-queue-multiple-servers/3218/2 "2022-07-20T12:42:07Z")

</div>

Update: I had removed the nodes and recreated with qmgr. Now the pbdsnodes -av gives me a single results with a free state:

sysadmin@Precision-7920-Tower:~/testVASP/newtest$ pbsnodes -av  
Precision-7920-Tower  
Mom = precision-7920-tower  
ntype = PBS  
state = free  
pcpus = 20  
resources\_available.arch = linux  
resources\_available.host = precision-7920-tower  
resources\_available.mem = 97495476kb  
resources\_available.ncpus = 20  
resources\_available.vnode = Precision-7920-Tower  
resources\_assigned.accelerator\_memory = 0kb  
resources\_assigned.hbmem = 0kb  
resources\_assigned.mem = 0kb  
resources\_assigned.naccelerators = 0  
resources\_assigned.ncpus = 0  
resources\_assigned.vmem = 0kb  
resv\_enable = True  
sharing = default\_shared  
license = l  
last\_state\_change\_time = Wed Jul 20 15:38:04 2022

This still does not solve the issue with the job in queue. I will wait for answers.

Thanks!

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [July 21, 2022, 6:37am UTC](https://community.openpbs.org/t/job-stuck-in-queue-multiple-servers/3218/3 "2022-07-21T06:37:10Z")

</div>

Please share us the output of the below commands:

- qstat -answ1
- qstat -fx 
- qstat -Bf

---

<div class="post-metadata">

**Author:** ![Blackest](https://avatars.discourse-cdn.com/v4/letter/b/5e9695/32.png) [@Blackest](https://community.openpbs.org/u/Blackest)\
**Post date:** [August 16, 2022, 10:37am UTC](https://community.openpbs.org/t/job-stuck-in-queue-multiple-servers/3218/4 "2022-08-16T10:37:18Z")

</div>

Sorry for the late answer. This are the outputs:  
sysadmin@Precision-7920-Tower:~/testVASP/newtest$ qstat -answ

Precision-7920-Tower:  
Req’d Req’d Elap  
Job ID Username Queue Jobname SessID NDS TSK Memory Time S Time

* * *

4011.Precision-7920-Tower sysadmin workq testjob – 1 20 – 24:00 Q –  
–  
–

sysadmin@Precision-7920-Tower:~/testVASP/newtest$ qstat -fx  
qstat: PBS is not configured to maintain job history

sysadmin@Precision-7920-Tower:~/testVASP/newtest$ qstat -Bf  
Server: Precision-7920-Tower  
server\_state = Active  
server\_host = precision-7920-tower  
scheduling = True  
total\_jobs = 1  
state\_count = Transit:0 Queued:1 Held:0 Waiting:0 Running:0 Exiting:0 Begun  
:0  
managers = root@Precision-7920-Tower  
default\_queue = workq  
log\_events = 511  
mailer = /usr/sbin/sendmail  
mail\_from = adm  
query\_other\_jobs = True  
resources\_default.ncpus = 1  
default\_chunk.ncpus = 1  
resources\_max.mpiprocs = 20  
resources\_max.ncpus = 20  
scheduler\_iteration = 600  
resv\_enable = True  
node\_fail\_requeue = 310  
max\_array\_size = 10000  
pbs\_license\_min = 0  
pbs\_license\_max = 2147483647  
pbs\_license\_linger\_time = 31536000  
license\_count = Avail\_Global:1000000 Avail\_Local:1000000 Used:0 High\_Use:0  
pbs\_version = 20.0.0  
eligible\_time\_enable = False  
max\_concurrent\_provision = 5  
max\_job\_sequence\_id = 9999999

Thanks!

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [August 17, 2022, 2:04pm UTC](https://community.openpbs.org/t/job-stuck-in-queue-multiple-servers/3218/5 "2022-08-17T14:04:45Z")

</div>

Please try this:

1. qmgr -c “print nodes @default” \> printnodes.txt
2. qmgr -c “delete nodes @default”
3. qmgr -c “create node precision-7920-tower”
4. Then submit a test job
5. qstat -answ1
6. pbsnodes -av

---

<div class="post-metadata">

**Author:** ![Blackest](https://avatars.discourse-cdn.com/v4/letter/b/5e9695/32.png) [@Blackest](https://community.openpbs.org/u/Blackest)\
**Post date:** [September 14, 2022, 4:08pm UTC](https://community.openpbs.org/t/job-stuck-in-queue-multiple-servers/3218/6 "2022-09-14T16:08:50Z")

</div>

Sorry for the delayed reply. Everything seems to work fine now and I am not clear what happened. I will let you know if I encounter other issues. Meanwhile we can consider the issue solved.

Thank you again!
