# My job stay queued

**URL:** <https://community.openpbs.org/t/my-job-stay-queued/1315>\
**Category:** Users/Site Administrators\
**Created:** [November 18, 2018, 3:14pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315 "2018-11-18T15:14:38Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![nekcorp](https://avatars.discourse-cdn.com/v4/letter/n/a587f6/32.png) [@nekcorp](https://community.openpbs.org/u/nekcorp)\
**Post date:** [November 18, 2018, 3:14pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/1 "2018-11-18T15:14:38Z")

</div>

Hi,

Since today and I don’t know why, but when I submit a job it is staying queued.

what should I do to understand what is the problem ?

Yesterday all runing perfectly.

Thank a lot for your help

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [November 18, 2018, 5:31pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/2 "2018-11-18T17:31:06Z")

</div>

Please share the output of the below commands

1. qstat -answ1
2. pbsnodes -av

Note;

- please check whether all the nodes status is free ? pbsnodes -av | grep -e Mom -e state
- please check whether the job requests can be matched on to the compute resources , otherwise, job will be in the queue ?

---

<div class="post-metadata">

**Author:** ![nekcorp](https://avatars.discourse-cdn.com/v4/letter/n/a587f6/32.png) [@nekcorp](https://community.openpbs.org/u/nekcorp)\
**Post date:** [November 21, 2018, 2:58pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/3 "2018-11-21T14:58:25Z")

</div>

$ qstat -answ1

return nothing

> $ pbsnodes -av

return :

centos7  
Mom = centos7.home  
ntype = PBS  
state = state-unknown,down  
pcpus = 1  
resources\_available.host = centos7  
resources\_available.ncpus = 1  
resources\_available.vnode = centos7  
resources\_assigned.accelerator\_memory = 0kb  
resources\_assigned.hbmem = 0kb  
resources\_assigned.mem = 0kb  
resources\_assigned.naccelerators = 0  
resources\_assigned.ncpus = 0  
resources\_assigned.vmem = 0kb  
resv\_enable = True  
sharing = default\_shared

> $ pbsnodes -av | grep -e Mom -e state

return

```
 Mom = centos7.home
 state = state-unknown,down

```

> please check whether the job requests can be matched on to the compute resources , otherwise, job will be in the queue ?

I submit jobs I have already submitted, so normally compute resources are ok.

I do not understand the problem.

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [November 21, 2018, 5:39pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/4 "2018-11-21T17:39:11Z")

</div>

> [@nekcorp](#):
>
> $ qstat -answ1
> 
> return nothing

Then there is no jobs in the queue. Hence, please submit sample sleep jobs as below  
qsub – /bin/sleep 100

> [@nekcorp](#):
>
> Mom = centos7.home state = state-unknown,down

The status of the compute node is down, hence job is still in the queue.

**Why the node is down** - communication issues between the PBS Server and PBS Mom (Compute Node)

1. pbs\_mom service might not be running on the compute node centos7.home
2. firewall might be blocking the ports ( 15001 - 15007 , 17001 ) between the headnode and compute node (vice versa) . Disable firewall completely and check .
3. DNS resolution ( forward and reverse resolution of the compute node / headnode ) from respective systems.
4. Check SELinux is disabled and system is rebooted after disabling the SELinux

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [December 7, 2018, 12:04pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/5 "2018-12-07T12:04:56Z")

</div>

Hi. I also have a problem. When I write the job submission script and specify a particular node name, the job stays in a queue. #PBS -l nodes=compunode-0-3.local

After submitting the job which stays in the queue, i use this command qstat -answ1 i get this error  
Can Never Run: Insufficient amount of resource: host (compunode-0-3.local !=compunode-0-1,compunode-0-2,compunode.-0-3,…

We have the followiing restrictions on the server for every user(PBS\_GENERIC)

max\_run=3  
max\_run\_res.ncpus=72  
max\_run\_res.nodect=2  
max\_queued=2

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [December 9, 2018, 10:31pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/6 "2018-12-09T22:31:47Z")

</div>

Please share the output of the below command  
pbsnodes computenode-0-3.local

Can you please try this command:  
qsub -l host=computenode-0-3 – /bin/sleep 100

> [@vincent718](#):
>
> Can Never Run: Insufficient amount of resource: host (compunode-0-3.local !=compunode-0-1,compunode-0-2,compunode.-0-3,…

It seems the “mom name” is not matching the request.

- mom name should have the short name , please check

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [December 10, 2018, 11:12am UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/7 "2018-12-10T11:12:54Z")

</div>

These are the results.

[user@login ~]# pbsnodes compute-0-3.local  
Node: compute-0-3.local, Error: Unknown node  
[user@login ~]# pbsnodes compute-0-3  
compute-0-3.local  
Mom = compute-0-3.local  
ntype = PBS  
state = free  
pcpus = 36  
resources\_available.arch = linux  
resources\_available.host = compute-0-3  
resources\_available.mem = 263727076kb  
resources\_available.ncpus = 36  
resources\_available.ngpus = 2  
resources\_available.vnode = compute-0-3.local  
resources\_assigned.accelerator\_memory = 0kb  
resources\_assigned.mem = 0kb  
resources\_assigned.naccelerators = 0  
resources\_assigned.ncpus = 0  
resources\_assigned.ngpus = 0  
resources\_assigned.vmem = 0kb  
resv\_enable = True  
sharing = default\_shared

[user@login ~]$ qsub -l host=compute-0-3 – /bin/sleep 100  
usage: qsub [-a date\_time] [-A account\_string] [-c interval]  
[-C directive\_prefix] [-e path] [-f] [-h] [-I [-X]] [-j oe|eo] [-J X-Y[:Z]]  
[-k keep] [-l resource\_list] [-m mail\_options] [-M user\_list]  
[-N jobname] [-o path] [-p priority] [-P project] [-q queue] [-r y|n]  
[-R o|e|oe] [-S path] [-u user\_list] [-W otherattributes=value…]  
[-S path] [-u user\_list] [-W otherattributes=value…]  
[-v variable\_list] [-V] [-z] [script | – command [arg1 …]]  
qsub --version

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [December 10, 2018, 5:25pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/8 "2018-12-10T17:25:17Z")

</div>

The command should be

qsub -l host=compute-0-3 - - /bin/sleep 100

qsub \< hyphen \>\< l for london \>\< space \>host=compute-0-3\< space \> \< hyphen \>\< hyphen \>\< space\> /bin/sleep 1000

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [December 11, 2018, 9:02am UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/9 "2018-12-11T09:02:04Z")

</div>

This is the command and output

qsub -l host=compute-0-3 – bin/sleep 100  
80126.master1.local

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [December 11, 2018, 3:45pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/10 "2018-12-11T15:45:49Z")

</div>

Thank you !

> [@vincent718](#):
>
> qsub -l host=compute-0-3 – /bin/sleep 100 # there was a forward / missing in /bin/sleep 100

- did the job run on the requested host ?  
please share the output of
- qstat -answ1
- qstat -fx 80126

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [December 11, 2018, 4:01pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/11 "2018-12-11T16:01:31Z")

</div>

Yes it did run. Here are the other outputs  
qstat -answ1

100363.master1.local user workq STDIN 17184 1 1 – – R 00:00:00 compute-0-3  
Job run at Tue Dec 11 at 15:58 on (compute-0-3.local:ncpus=1)

qstat -fx 100363  
Job Id: 100363.master1.local  
Job\_Name = STDIN  
Job\_Owner = user@login.local  
resources\_used.cpupercent = 0  
resources\_used.cput = 00:00:00  
resources\_used.mem = 348kb  
resources\_used.ncpus = 1  
resources\_used.vmem = 4316kb  
resources\_used.walltime = 00:01:40  
job\_state = F  
queue = workq  
server = master1.local  
Checkpoint = u  
ctime = Tue Dec 11 15:58:25 2018  
Error\_Path = login.local:/home/user/STDIN.e100363  
exec\_host = compute-0-3.local/0  
exec\_vnode = (compute-0-3.local:ncpus=1)  
Hold\_Types = n  
Join\_Path = n  
Keep\_Files = n  
Mail\_Points = a  
mtime = Tue Dec 11 16:00:06 2018  
Output\_Path = login.local:/home/user/STDIN.o100363  
Priority = 0  
qtime = Tue Dec 11 15:58:25 2018  
Rerunable = True  
Resource\_List.host = compute-0-3  
Resource\_List.ncpus = 1  
Resource\_List.nodect = 1  
Resource\_List.place = pack  
Resource\_List.select = 1:host=compute-0-3:ncpus=1  
stime = Tue Dec 11 15:58:25 2018  
session\_id = 17184  
jobdir = /home/user  
substate = 92  
Variable\_List = PBS\_O\_HOME=/home/user,PBS\_O\_LANG=en\_US.UTF-8,  
PBS\_O\_LOGNAME=user-l host=compute-0-3:ncpus=10  
,  
PBS\_O\_PATH=/opt/apps/intel/compilers\_and\_libraries\_2017.4.196/linux/mp  
i/intel64/bin:/opt/apps/intel/compilers\_and\_libraries\_2017.4.196/linux/  
bin/intel64:/usr/local/bin:/usr/bin:/usr/local/sbin:/usr/sbin:/opt/ibut  
ils/bin:/opt/pbs/bin:/home/user/.local/bin:/home/user/bin,  
PBS\_O\_MAIL=/var/spool/mail/user,PBS\_O\_SHELL=/bin/bash,  
PBS\_O\_WORKDIR=/home/user,PBS\_O\_SYSTEM=Linux,PBS\_O\_QUEUE=workq,  
PBS\_O\_HOST=login.local  
comment = Job run at Tue Dec 11 at 15:58 on (compute-0-3.local:ncpus=1) and  
finished  
etime = Tue Dec 11 15:58:25 2018  
run\_count = 1  
Stageout\_status = 1  
Exit\_status = 0  
Submit\_arguments = -l host=compute-0-3 – /bin/sleep 100  
executable = jsdl-hpcpa:Executable/bin/sleep\</jsdl-hpcpa:Executable\>  
argument\_list = jsdl-hpcpa:Argument100\</jsdl-hpcpa:Argument\>  
history\_timestamp = 1544544006  
project = \_pbs\_project\_default

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [December 12, 2018, 5:47am UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/12 "2018-12-12T05:47:05Z")

</div>

Thank you. It is all working now.  
Do you still see any issues ?

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [December 12, 2018, 7:04am UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/13 "2018-12-12T07:04:39Z")

</div>

But what is the syntax for selecting nodes in PBS?  
I am using pbs pro version 17 and i get an error when i use this command below  
#PBS -l host=compute-0-3:ncpus=10 -l mem=10GB  
The error is  
Illegal attribute or resource value Resource\_List.select

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [December 12, 2018, 7:22am UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/14 "2018-12-12T07:22:26Z")

</div>

I got it now . The correct syntax is

#PBS -l host=compute-0-3 -l ncpus=10 -l mem=2GB

Thank you very much adarsh.

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [December 13, 2018, 3:48pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/15 "2018-12-13T15:48:52Z")

</div>

I have detected an issue.

When i enter the host name in job submission script, the job is submitted to a different host( .  
eg. if i select compute node 5 #PBS -l host=compute-0-5 -l ncpus=18 -l mem=32GB

The job gets submitted to a compute node which is free ( starting from 1, 2, or 3 or higher).

But when i select the node using the interactive option such as  
qsub -I -l host=compute-0-5 -l ncpus=10 -l mem=2GB

This rather works

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [December 14, 2018, 12:48pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/16 "2018-12-14T12:48:46Z")

</div>

**Submit a job to a particular host:**

`qsub -l select=1:ncpus=2:mem=32gb:host=compute-0-5 -- /bin/sleep 100`

**Submit a job and let PBS decide where to run:**

`qsub -l select=1:ncpus=2:mem=32gb -- /bin/sleep 100`

**To run 2 cpu jobs on two nodes requesting 16GB on each of the nodes**

`qsub -l select=2:ncpus=1:mem=16gb -l place=scatter -- /bin/sleep 100`

**To specifically run jobs on two specific hosts**

`qsub -l nodes=compute-0-3+compute-0-5 -- /bin/sleep 100`

`qsub -l select=1:ncpus=1:mem=10gb:host=compute-0-3+1:mem=10gb:host=compute-0-5 -- /bin/sleep 100`

Please read the below admin guide section: **5.4.9 Job-wide vs. Chunk Resources**

> **[PBS18.2\_BigBook.pdf](https://www.pbsworks.com/pdfs/PBS18.2_BigBook.pdf)**
>
> 24.13 MB

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [January 24, 2020, 5:39pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/17 "2020-01-24T17:39:09Z")

</div>

Hi Adash  
I did a restart of our HPC system and now all jobs are in a queue.  
I tried the following command ‘pdsh date’ and i got the values below

ompute-0-2: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-2: @ WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED! @  
compute-0-2: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-4: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-2: IT IS POSSIBLE THAT SOMEONE IS DOING SOMETHING NASTY!  
compute-0-1: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-4: @ WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED! @  
compute-0-2: Someone could be eavesdropping on you right now (man-in-the-middle attack)!  
compute-0-3: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-1: @ WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED! @  
compute-0-4: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-2: It is also possible that a host key has just been changed.  
compute-0-2: The fingerprint for the ECDSA key sent by the remote host is  
compute-0-1: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-4: IT IS POSSIBLE THAT SOMEONE IS DOING SOMETHING NASTY!  
compute-0-3: @ WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED! @  
compute-0-4: Someone could be eavesdropping on you right now (man-in-the-middle attack)!  
compute-0-1: IT IS POSSIBLE THAT SOMEONE IS DOING SOMETHING NASTY!  
compute-0-2: SHA256:gFtWek5zWZJQCnNQwycLTUtZXiZudM5S9H+xkK860ok.  
compute-0-3: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-4: It is also possible that a host key has just been changed.  
compute-0-1: Someone could be eavesdropping on you right now (man-in-the-middle attack)!  
compute-0-2: Please contact your system administrator.  
compute-0-4: The fingerprint for the ECDSA key sent by the remote host is  
compute-0-1: It is also possible that a host key has just been changed.  
compute-0-3: IT IS POSSIBLE THAT SOMEONE IS DOING SOMETHING NASTY!  
compute-0-2: Add correct host key in /root/.ssh/known\_hosts to get rid of this message.  
compute-0-1: The fingerprint for the ECDSA key sent by the remote host is  
compute-0-4: SHA256:o+nbSQpNh7f7YyQwP3myiY9H0LtKhwHWLBTlkXfggkE.  
compute-0-3: Someone could be eavesdropping on you right now (man-in-the-middle attack)!  
compute-0-2: Offending ECDSA key in /root/.ssh/known\_hosts:3  
compute-0-1: SHA256:1XCBEBwL4CIsAB+XU1uGM8borPm6WR1p+V1isuiNRFE.  
compute-0-4: Please contact your system administrator.  
compute-0-2: Password authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-3: It is also possible that a host key has just been changed.  
compute-0-1: Please contact your system administrator.  
compute-0-4: Add correct host key in /root/.ssh/known\_hosts to get rid of this message.  
compute-0-2: Keyboard-interactive authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-3: The fingerprint for the ECDSA key sent by the remote host is  
compute-0-1: Add correct host key in /root/.ssh/known\_hosts to get rid of this message.  
compute-0-4: Offending ECDSA key in /root/.ssh/known\_hosts:5  
compute-0-3: SHA256:mq67n0i6+BiZPVWWeHt5iORDaYXIiqsrhvzYQjr2YkQ.  
compute-0-1: Offending ECDSA key in /root/.ssh/known\_hosts:1  
compute-0-4: Password authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-3: Please contact your system administrator.  
compute-0-1: Password authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-4: Keyboard-interactive authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-3: Add correct host key in /root/.ssh/known\_hosts to get rid of this message.  
compute-0-1: Keyboard-interactive authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-3: Offending ECDSA key in /root/.ssh/known\_hosts:4  
compute-0-3: Password authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-3: Keyboard-interactive authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-5: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-5: @ WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED! @  
compute-0-5: @@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@  
compute-0-5: IT IS POSSIBLE THAT SOMEONE IS DOING SOMETHING NASTY!  
compute-0-5: Someone could be eavesdropping on you right now (man-in-the-middle attack)!  
compute-0-5: It is also possible that a host key has just been changed.  
compute-0-5: The fingerprint for the ECDSA key sent by the remote host is  
compute-0-5: SHA256:Ziq6yZmX0BRFHNwCD5U5425ef0B6ja4q+Q5h1Exs12U.  
compute-0-5: Please contact your system administrator.  
compute-0-5: Add correct host key in /root/.ssh/known\_hosts to get rid of this message.  
compute-0-5: Offending ECDSA key in /root/.ssh/known\_hosts:6  
compute-0-5: Password authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-5: Keyboard-interactive authentication is disabled to avoid man-in-the-middle attacks.  
compute-0-2: Fri Jan 24 17:38:38 UTC 2020  
compute-0-3: Fri Jan 24 17:38:38 UTC 2020  
compute-0-1: Fri Jan 24 17:38:39 UTC 2020  
compute-0-4: Fri Jan 24 17:38:37 UTC 2020  
compute-0-5: Fri Jan 24 17:38:38 UTC 2020  
compute-0-6: ssh: connect to host compute-0-6.local port 22: No route to host  
pdsh@master: compute-0-6: ssh exited with exit code 255

---

<div class="post-metadata">

**Author:** ![mkaro](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.openpbs.org/mkaro/32/85_2.png) [@mkaro](https://community.openpbs.org/u/mkaro)\
**Post date:** [January 24, 2020, 6:24pm UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/18 "2020-01-24T18:24:15Z")

</div>

I see two separate problems. One has to do with your ssh configuration…

> [@vincent718](#):
>
> compute-0-5: Offending ECDSA key in /root/.ssh/known\_hosts:6

You should not run jobs as the root user. Also, it appears something has changed (IP address?) to invalidate the existing ssh keys.

The second problem is this…

> [@vincent718](#):
>
> compute-0-6: ssh: connect to host compute-0-6.local port 22: No route to host

This indicates there is something wrong with your network configuration.

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [January 25, 2020, 3:00am UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/19 "2020-01-25T03:00:15Z")

</div>

Hi Michael,

What do you suggest I do? Everything was working perfectly until I did a restart of the server. By the way we have two master nodes and one of them works as a fail over( depends on which is available or not)

Vincent Appiah

![](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.pbspro.org/mkaro/45/85_2.png "mkaro")

[mkaro](http://community.openpbs.org/u/mkaro)

```
    January 24

```

I see two separate problems. One has to do with your ssh configuration…

> ![](https://avatars.discourse.org/v4/letter/v/f9ae1b/40.png) vincent718:  
> compute-0-5: Offending ECDSA key in /root/.ssh/known\_hosts:6

You should not run jobs as the root user. Also, it appears something has changed (IP address?) to invalidate the existing ssh keys.

The second problem is this…

> ![](https://avatars.discourse.org/v4/letter/v/f9ae1b/40.png) vincent718:  
> compute-0-6: ssh: connect to host compute-0-6.local port 22: No route to host

This indicates there is something wrong with your network configuration.

---

<div class="post-metadata">

**Author:** ![vincent718](https://avatars.discourse-cdn.com/v4/letter/v/f9ae1b/32.png) [@vincent718](https://community.openpbs.org/u/vincent718)\
**Post date:** [January 25, 2020, 9:33am UTC](https://community.openpbs.org/t/my-job-stay-queued/1315/20 "2020-01-25T09:33:34Z")

</div>

Hello Michael,

Before I performed the restart, i used _qdel $(select)_ to delete all existing jobs. This was done using the root account. Could that be the cause?

Hi Michael,

What do you suggest I do? Everything was working perfectly until I did a restart of the server. By the way we have two master nodes and one of them works as a fail over( depends on which is available or not)

Vincent Appiah

![](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.pbspro.org/mkaro/45/85_2.png "mkaro")

[mkaro](http://community.openpbs.org/u/mkaro)

```
    January 24

```

I see two separate problems. One has to do with your ssh configuration…

> ![](https://avatars.discourse.org/v4/letter/v/f9ae1b/40.png) vincent718:  
> compute-0-5: Offending ECDSA key in /root/.ssh/known\_hosts:6

You should not run jobs as the root user. Also, it appears something has changed (IP address?) to invalidate the existing ssh keys.

The second problem is this…

> ![](https://avatars.discourse.org/v4/letter/v/f9ae1b/40.png) vincent718:  
> compute-0-6: ssh: connect to host compute-0-6.local port 22: No route to host

This indicates there is something wrong with your network configuration.

[Next page](https://community.openpbs.org/t/my-job-stay-queued/1315.md?page=2)
