# Test job waits in queue

**URL:** <https://community.openpbs.org/t/test-job-waits-in-queue/3019>\
**Category:** Users/Site Administrators\
**Created:** [March 1, 2022, 9:05am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019 "2022-03-01T09:05:13Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![watzinki](https://avatars.discourse-cdn.com/v4/letter/w/f04885/32.png) [@watzinki](https://community.openpbs.org/u/watzinki)\
**Post date:** [March 1, 2022, 9:05am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/1 "2022-03-01T09:05:13Z")

</div>

Hi All,  
I am new to pbspro. I installed PBS server and execute rpm packages on the master and the slave node respectively. I tried to set up everything following instructions on the web, however, when I run a test job it waits in the queue forever. Any help would be appreciated.

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [March 1, 2022, 11:04am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/2 "2022-03-01T11:04:36Z")

</div>

Please share the output of the below command:

> 1. qstat -answ1
> 2. pbsnodes -av
> 3. qmgr -c “print server”

---

<div class="post-metadata">

**Author:** ![watzinki](https://avatars.discourse-cdn.com/v4/letter/w/f04885/32.png) [@watzinki](https://community.openpbs.org/u/watzinki)\
**Post date:** [March 1, 2022, 3:14pm UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/3 "2022-03-01T15:14:16Z")

</div>

I actually have one master node called hep-node0 and a slave node. I actually set up the master node in a way so that I could use it for job execution as well.

Below is the output of the commands I executed as directed.

[ali\_0@hep-node0 ~]$ qstat -answ1  
hep-node0:  
Req’d Req’d Elap  
Job ID Username Queue Jobname SessID NDS TSK Memory Time S Time

* * *

1014.hep-node0 ali\_0 batch example-job.sh – 1 1 – 00:00 Q – –  
Not Running: Not enough free nodes available  
1015.hep-node0 ali\_0 batch STDIN – 1 1 – 00:00 Q – –  
Not Running: Not enough free nodes available  
[ali\_0@hep-node0 ~]$ pbsnodes -av  
ali\_2  
Mom = hep-node2  
ntype = PBS  
state = state-unknown,down  
pcpus = 1  
resources\_available.host = hep-node2  
resources\_available.ncpus = 1  
resources\_available.vnode = ali\_2  
resources\_assigned.accelerator\_memory = 0kb  
resources\_assigned.hbmem = 0kb  
resources\_assigned.mem = 0kb  
resources\_assigned.naccelerators = 0  
resources\_assigned.ncpus = 0  
resources\_assigned.vmem = 0kb  
comment = node down: communication closed  
resv\_enable = True  
sharing = default\_shared  
last\_state\_change\_time = Mon Feb 28 23:02:12 2022

[ali\_0@hep-node0 ~]$ qmgr -c “print server”  
Unknown Host.  
qmgr: cannot connect to server server”

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [March 1, 2022, 3:53pm UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/4 "2022-03-01T15:53:26Z")

</div>

> [@watzinki](#):
>
> state = state-unknown,down

Please check the output of pbsnodes -av

1. node status is down reason might be

- The pbs\_mom service is down ( systemctl status pbs # on the compute node ) or ps -ef | grep pbs\_mom
- hence your job is in the queued status , there aren’t enough resources available to run your job.

Make sure

- pbs\_mom service is up and running
- cat /etc/pbs.conf | grep PBS\_START\_MOM  
PBS\_START\_MOM=1

> [@watzinki](#):
>
> [ali\_0@hep-node0 ~]$ qmgr -c “print server”  
> Unknown Host.  
> qmgr: cannot connect to server server”

Nothing wrong with this , you need to type the command instead of copy pasting it.  
Copy pasting it has some special character and hence it fails.

---

<div class="post-metadata">

**Author:** ![watzinki](https://avatars.discourse-cdn.com/v4/letter/w/f04885/32.png) [@watzinki](https://community.openpbs.org/u/watzinki)\
**Post date:** [March 1, 2022, 5:38pm UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/5 "2022-03-01T17:38:28Z")

</div>

> [@watzinki](#):
>
> qmgr -c “print server”

Thank you for your replies. pbs\_mom service seems to be up and running on the master node “hep-node0”. Below is the output of “systemctl status PBS” command.  
[ali\_0@hep-node0 ~]$ systemctl status pbs  
● pbs.service - Portable Batch System  
Loaded: loaded (/opt/pbs/libexec/pbs\_init.d; enabled; vendor preset: disabled)  
Active: active (running) since Mon 2022-02-28 23:02:07 EET; 1 day 2h ago  
Docs: man:pbs(8)  
Process: 1245 ExecStart=/opt/pbs/libexec/pbs\_init.d start (code=exited, status=0/SUCCESS)  
Tasks: 14  
Memory: 17.1M  
CGroup: /system.slice/pbs.service  
├─1412 /opt/pbs/sbin/pbs\_comm  
├─1466 /opt/pbs/sbin/pbs\_mom  
├─1516 /opt/pbs/sbin/pbs\_sched  
├─1877 /opt/pbs/sbin/pbs\_ds\_monitor monitor  
├─1921 /usr/bin/postgres -D /var/spool/pbs/datastore -p 15007  
├─1929 postgres: logger process  
├─1931 postgres: checkpointer process  
├─1932 postgres: writer process  
├─1933 postgres: wal writer process  
├─1934 postgres: autovacuum launcher process  
├─1935 postgres: stats collector process  
├─2075 postgres: postgres pbs\_datastore 192.168.1.1(59718) idle  
└─2076 /opt/pbs/sbin/pbs\_server.bin

Feb 28 23:02:06 [hep-node0.com](http://hep-node0.com) systemd[1]: Starting Portable Batch System…  
Feb 28 23:02:07 [hep-node0.com](http://hep-node0.com) systemd[1]: Started Portable Batch System.  
Feb 28 23:02:07 [hep-node0.com](http://hep-node0.com) su[1610]: (to postgres) root on none  
Feb 28 23:02:07 [hep-node0.com](http://hep-node0.com) su[1687]: (to postgres) root on none  
Feb 28 23:02:07 [hep-node0.com](http://hep-node0.com) su[1754]: (to postgres) root on none  
Feb 28 23:02:07 [hep-node0.com](http://hep-node0.com) su[1792]: (to postgres) root on none  
Feb 28 23:02:07 [hep-node0.com](http://hep-node0.com) su[1878]: (to postgres) root on none  
Feb 28 23:02:12 [hep-node0.com](http://hep-node0.com) pbs\_init.d[1245]: Starting PBS in background

**[ali\_0@hep-node0 ~]$ qmgr -c “print server”**

# 

# Create queues and set their attributes.

# 

# 

# Create and define queue batch

# 

create queue batch  
set queue batch queue\_type = Execution  
set queue batch resources\_default.ncpus = 1  
set queue batch resources\_default.nodect = 1  
set queue batch resources\_default.nodes = 1  
set queue batch resources\_default.walltime = 00:00:36  
set queue batch enabled = True  
set queue batch started = True

# 

# Set server attributes.

# 

set server scheduling = True  
set server acl\_roots = username@\*  
set server operators = username@\*  
set server default\_queue = batch  
set server log\_events = 511  
set server mail\_from = adm  
set server query\_other\_jobs = True  
set server resources\_default.ncpus = 1  
set server default\_chunk.ncpus = 1  
set server scheduler\_iteration = 600  
set server flatuid = True  
set server resv\_enable = True  
set server node\_fail\_requeue = 310  
set server max\_array\_size = 10000  
set server pbs\_license\_min = 0  
set server pbs\_license\_max = 2147483647  
set server pbs\_license\_linger\_time = 31536000  
set server eligible\_time\_enable = False  
set server max\_concurrent\_provision = 5  
set server max\_job\_sequence\_id = 9999999

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [March 1, 2022, 9:47pm UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/6 "2022-03-01T21:47:05Z")

</div>

Thank you for the above information

> [@watzinki](#):
>
> resources\_available.vnode = ali\_2

This should be also hep-node2 , not sure why it is ali\_2.

```auto
For example, a similar setup like yours, would have this output:
[root@demo ~]# pbsnodes -av | grep demo
demo
     Mom = demo
     resources_available.host = demo
     resources_available.vnode = demo
[root@demo ~]# cat /etc/hosts | grep demo
192.168.64.128 demo
[root@demo ~]# cat /etc/pbs.conf | grep MOM
PBS_START_MOM=1
[root@demo ~]# cat /var/spool/pbs/mom_priv/config | grep client
$clienthost demo
[root@demo ~]# ps -ef | grep pbs_mom
root 1620 1 0 Feb28 ? 00:00:07 /opt/pbs/sbin/pbs_mom
root 1129595 1126933 0 21:45 pts/0 00:00:00 grep --color=auto pbs_mom

```

Please make sure

1. selinux is disabled and system is rebooted
2. ports 15001 to 15009 and 17001 is open for communication (or firewall / ip tables completely disabled)
3. static IP address and hostname and /etc/hosts is up-to-date / DNS resolvable

---

<div class="post-metadata">

**Author:** ![watzinki](https://avatars.discourse-cdn.com/v4/letter/w/f04885/32.png) [@watzinki](https://community.openpbs.org/u/watzinki)\
**Post date:** [March 2, 2022, 7:14am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/7 "2022-03-02T07:14:58Z")

</div>

Dear @adarsh , thanks again for your answers. I got the master node running. I can run the jobs on it, but cannot run on the slave node, which is the ali\_2 machine or whose hostname is “hep-node2”.

selinux already disabled and rebooted on both machines. I established ssh connection between two machines with passwordless access to each other as well.  
With “systemctl status pbs” on the slave node, I see that PBS running on the node. However, when I tried to set the ali\_2 machine as slave node via the command qmgr -c “create node hep-node2”, I got the error message “No route to host  
qmgr: cannot connect to server” .So, my naive first thought was that this is due to the requirement for the ports to be opened. If you think the way like me, could you please elaborate a bit on how to get these ports opened on centos 7? Also, do these ports have to be open on both slave and master nodes or on the slave only?

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [March 2, 2022, 8:10am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/8 "2022-03-02T08:10:54Z")

</div>

> [@watzinki](#):
>
> However, when I tried to set the ali\_2 machine as slave node via the command qmgr -c “create node hep-node2”, I got the error message “No route to host  
> qmgr: cannot connect to server”

Please check

1. DNS / Static IP / hostname / check /etc/hosts ( network address resolution is important)
2. systemctl stop firewalld ; systemctl disable firewalld  
#otherwise add the above mentioned ports to be allowed in firewalld  
firewall-cmd --zone=public --permanent --add-port=15001/tcp # for all ports  
firewall-cmd --reload

> [@watzinki](#):
>
> could you please elaborate a bit on how to get these ports opened on centos 7?

```
   firewall-cmd --zone=public --permanent --add-port=15001/tcp # for all ports
   firewall-cmd --reload

```

> [@watzinki](#):
>
> Also, do these ports have to be open on both slave and master nodes or on the slave only?

Yes

If you are interested, please refer: [Building a PBS Professional Virtual Test Cluster with Ubuntu](https://www.altair.com/resource/building-a-pbs-professional-virtual-test-cluster-with-ubuntu)

---

<div class="post-metadata">

**Author:** ![watzinki](https://avatars.discourse-cdn.com/v4/letter/w/f04885/32.png) [@watzinki](https://community.openpbs.org/u/watzinki)\
**Post date:** [March 2, 2022, 11:09am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/9 "2022-03-02T11:09:36Z")

</div>

Thank you again!

Seems the master node sees the slave or vice versa.  
After firewall settings, pbsnodes -av output on both master node and slave node as the following.  
One more question. Is there a way to check if the submitted jobs are executed on the slave node as well? Although it seems nodes can contact each other, I doubt, the test job is only running on the master node.

hep-node0  
Mom = hep-node0  
ntype = PBS  
state = free  
pcpus = 4  
resources\_available.arch = linux  
resources\_available.host = hep-node0  
resources\_available.mem = 16265032kb  
resources\_available.ncpus = 4  
resources\_available.vnode = hep-node0  
resources\_assigned.accelerator\_memory = 0kb  
resources\_assigned.hbmem = 0kb  
resources\_assigned.mem = 0kb  
resources\_assigned.naccelerators = 0  
resources\_assigned.ncpus = 0  
resources\_assigned.vmem = 0kb  
resv\_enable = True  
sharing = default\_shared  
last\_state\_change\_time = Wed Mar 2 12:08:35 2022  
last\_used\_time = Wed Mar 2 12:20:41 2022

hep-node2  
Mom = hep-node2  
ntype = PBS  
state = free  
pcpus = 1  
resources\_available.arch = linux  
resources\_available.host = hep-node2  
resources\_available.mem = 8007520kb  
resources\_available.ncpus = 4  
resources\_available.vnode = hep-node2  
resources\_assigned.accelerator\_memory = 0kb  
resources\_assigned.hbmem = 0kb  
resources\_assigned.mem = 0kb  
resources\_assigned.naccelerators = 0  
resources\_assigned.ncpus = 0  
resources\_assigned.vmem = 0kb  
resv\_enable = True  
sharing = default\_shared  
last\_state\_change\_time = Wed Mar 2 12:08:35 2022

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [March 3, 2022, 7:51am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/10 "2022-03-03T07:51:42Z")

</div>

Nice one !

> [@watzinki](#):
>
> One more question. Is there a way to check if the submitted jobs are executed on the slave node as well?

Try this :

```auto
submit as below : 
qsub -l select=1:ncpus=1 -l place=excl -- /bin/sleep 1000
qsub -l select=1:ncpus=1 -l place=excl -- /bin/sleep 1000
qstat -answ1 # this command should show you on which node(s) the job is running
ssh <node> 
ps -ef | grep sleep

```

---

<div class="post-metadata">

**Author:** ![watzinki](https://avatars.discourse-cdn.com/v4/letter/w/f04885/32.png) [@watzinki](https://community.openpbs.org/u/watzinki)\
**Post date:** [March 3, 2022, 9:02am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/11 "2022-03-03T09:02:49Z")

</div>

@adarsh I really appreciate your help and patience thus far. As I thought in the previous message, jobs are not sent to slave nodes but run on the master node only instead. qstat -answ1 command outputs the following.  
Job ID Username Queue Jobname SessID NDS TSK Memory Time S Time

* * *

2047.hep-node0 ali\_0 batch STDIN 5732 1 1 – 200:0 R 00:02 hep-node0/0  
Job run at Thu Mar 03 at 17:05 on (hep-node0:ncpus=1)  
2048.hep-node0 ali\_0 batch STDIN – 1 1 – 200:0 H – –  
job held, too many failed attempts to run

Also, I had opened the ports you mentioned above using the commands as you directed. However, some of the ports still seem to be closed. This might be the reason?

[ali\_0@hep-node0 ~]$ sudo nmap -p 15001-15009 192.168.1.1  
[sudo] password for ali\_0:

Starting Nmap 6.40 ( [http://nmap.org](http://nmap.org) ) at 2022-03-03 17:17 EET  
Nmap scan report for hep-node0 (192.168.1.1)  
Host is up (0.00012s latency).  
PORT STATE SERVICE  
15001/tcp open unknown  
15002/tcp open unknown  
15003/tcp open unknown  
15004/tcp open unknown  
15005/tcp closed unknown  
15006/tcp closed unknown  
15007/tcp open unknown  
15008/tcp closed unknown  
15009/tcp closed unknown

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [March 3, 2022, 4:13pm UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/12 "2022-03-03T16:13:05Z")

</div>

Thank you @watzinki , pleasure !

> [@watzinki](#):
>
> 2048.hep-node0 ali\_0 batch STDIN – 1 1 – 200:0 H – –  
> job held, too many failed attempts to run

Please check the mom logs of the second node  
source /etc/pbs.conf ; cd $PBS\_HOME/mom\_logs and check the logs for the day (YYYYMMDD)

Suspect on the second node:

1. user account does not exist or passwd not set
2. home directory missing
3. permission issues
4. The mom logs should be able to tell you the issue

---

<div class="post-metadata">

**Author:** ![watzinki](https://avatars.discourse-cdn.com/v4/letter/w/f04885/32.png) [@watzinki](https://community.openpbs.org/u/watzinki)\
**Post date:** [March 6, 2022, 9:08pm UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/13 "2022-03-06T21:08:42Z")

</div>

Dear @adarsh, I have been trying to figure out what causing the issue by myself for the last few days but no luck thus far. I set passwordless login among the nodes. I can test it by sshing from one node to another, back and forth without entering the password. However, It still seems no job is executed on the slave node but master node only. Below is the dump of the log file of the slave node called hep-node2. Hope you might have other suggestions.  
Thanks in advance.

03/06/2022 21:00:12;0100;pbs\_mom;Req;;Type 1 request received from root@192.168.1.1:15001, sock=1  
03/06/2022 21:00:12;0100;pbs\_mom;Req;;Type 5 request received from root@192.168.1.1:15001, sock=1  
03/06/2022 21:00:12;0028;pbs\_mom;Job;2067.hep-node0;No Password Entry for User ali\_0  
03/06/2022 21:00:12;0008;pbs\_mom;Job;2067.hep-node0;kill\_job  
03/06/2022 21:00:12;0100;pbs\_mom;Job;2067.hep-node0;hep-node2 cput=00:00:00 mem=0kb  
03/06/2022 21:00:12;0100;pbs\_mom;Job;2067.hep-node0;Obit sent  
03/06/2022 21:00:12;0100;pbs\_mom;Req;;Type 6 request received from root@192.168.1.1:15001, sock=1  
03/06/2022 21:00:12;0080;pbs\_mom;Job;2067.hep-node0;delete job request received  
03/06/2022 21:00:12;0008;pbs\_mom;Job;2067.hep-node0;kill\_job  
03/06/2022 21:00:12;0100;pbs\_mom;Req;;Type 1 request received from root@192.168.1.1:15001, sock=1  
03/06/2022 21:00:12;0100;pbs\_mom;Req;;Type 5 request received from root@192.168.1.1:15001, sock=1  
03/06/2022 21:00:12;0028;pbs\_mom;Job;2067.hep-node0;No Password Entry for User ali\_0  
03/06/2022 21:00:12;0008;pbs\_mom;Job;2067.hep-node0;kill\_job  
03/06/2022 21:00:12;0100;pbs\_mom;Job;2067.hep-node0;hep-node2 cput=00:00:00 mem=0kb  
03/06/2022 21:00:12;0100;pbs\_mom;Job;2067.hep-node0;Obit sent  
03/06/2022 21:00:12;0100;pbs\_mom;Req;;Type 6 request received from root@192.168.1.1:15001, sock=1  
03/06/2022 21:00:12;0080;pbs\_mom;Job;2067.hep-node0;delete job request received  
03/06/2022 21:00:12;0008;pbs\_mom;Job;2067.hep-node0;kill\_job  
03/06/2022 21:00:12;0100;pbs\_mom;Req;;Type 1 request received from root@192.168.1.1:15001, sock=1  
03/06/2022 21:00:12;0100;pbs\_mom;Req;;Type 5 request received from root@192.168.1.1:15001, sock=1  
03/06/2022 21:00:12;0028;pbs\_mom;Job;2067.hep-node0;No Password Entry for User ali\_0  
03/06/2022 21:00:12;0008;pbs\_mom;Job;2067.hep-node0;kill\_job

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [March 7, 2022, 7:52am UTC](https://community.openpbs.org/t/test-job-waits-in-queue/3019/14 "2022-03-07T07:52:31Z")

</div>

> [@watzinki](#):
>
> 03/06/2022 21:00:12;0028;pbs\_mom;Job;2067.hep-node0;No Password Entry for User ali\_0

This is the issue. It seems there is issue with user account and password or password is not set.
