# Execution node down

**URL:** <https://community.openpbs.org/t/execution-node-down/1723>\
**Category:** Users/Site Administrators\
**Created:** [August 6, 2019, 8:33pm UTC](https://community.openpbs.org/t/execution-node-down/1723 "2019-08-06T20:33:41Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![brownwrap](https://avatars.discourse-cdn.com/v4/letter/b/94ad74/32.png) [@brownwrap](https://community.openpbs.org/u/brownwrap)\
**Post date:** [August 6, 2019, 8:33pm UTC](https://community.openpbs.org/t/execution-node-down/1723/1 "2019-08-06T20:33:41Z")

</div>

Yesterday a user reported his jobs were going into the “H” state shortly after submitting the job. He said the last time this happened, a node was offlined. I did a tracejob on the job and it had attempted to run the 21 time and gave up. Also said ‘execution node down’. pbsnodes -l did not report that. I looked at the history and compute-0-15 had been put online. I offlined it and the job ran. Without the history being there, how would I determine the correct compute node. PBS must know. Thanks.

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [August 6, 2019, 8:52pm UTC](https://community.openpbs.org/t/execution-node-down/1723/2 "2019-08-06T20:52:24Z")

</div>

1. If the job has been attempted to run 21 times and gave up means one of these issues

2. Enable job history by setting qmgr -c ‘set server job\_history\_enable=true’  
Submit couple of jobs and check whether they go into H state , if yes, qstat -fx or tracejob , find out the compute node it was scheduled on, login to the compute nodes, check the mom logs for details

Some of the generic reasons for the job in “H” state

```
	* If there is an issue with authentication of the user on the compute node
	* or user home directory not mounted or home directory of the user not available on the compute nodes
	* not sure about the users authentication PBS keeps the job in held state
	* when the job is manually put on the hold state using qhold command
	* If the job is a dependent job

```

Caveats:  
\* does user has any issues logging onto the compute nodes?  
\* Can the user log in to the node?  
\* Is everything in order for the user account, username, password, home directory etc.

---

<div class="post-metadata">

**Author:** ![brownwrap](https://avatars.discourse-cdn.com/v4/letter/b/94ad74/32.png) [@brownwrap](https://community.openpbs.org/u/brownwrap)\
**Post date:** [August 6, 2019, 9:10pm UTC](https://community.openpbs.org/t/execution-node-down/1723/3 "2019-08-06T21:10:03Z")

</div>

I have history enabled. The mom\_logs don’t show anything. This happened yesterday and there is no log for that day:

-rw-r–r-- 1 root root 352 Aug 4 09:40 20190801  
drwxr-xr-x 2 root root 12K Aug 4 09:40 .  
-rw-r–r-- 1 root root 540 Aug 4 09:40 20190804  
[root@compute-0-15 mom\_logs]#

My question is, if tracejob shows “execution node down”, why doesn’t it tell me the node? I just guessed at the culprit, used pbsnodes to offline it and reran the job.

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [August 6, 2019, 9:15pm UTC](https://community.openpbs.org/t/execution-node-down/1723/4 "2019-08-06T21:15:27Z")

</div>

To offline the node, please use qmgr -c “set node NODENAME state=offline”.  
Do not offline the node using pbsnodes -o .  
#qmgr -c “set node NODENAME state=free” , to make is available/free.

Could you please share the output of qstat -fx , tracejob and pbs\_dtj , accounting log snippet of   
jobid – job id of the job which was put on H state

---

<div class="post-metadata">

**Author:** ![brownwrap](https://avatars.discourse-cdn.com/v4/letter/b/94ad74/32.png) [@brownwrap](https://community.openpbs.org/u/brownwrap)\
**Post date:** [August 6, 2019, 9:35pm UTC](https://community.openpbs.org/t/execution-node-down/1723/5 "2019-08-06T21:35:43Z")

</div>

Below is the output of tracejob. Once I offlined the node, ?I used qrun to run the job:

08/05/2019 16:47:09 A user=reinecke group=reinecke project=\_pbs\_project\_default jobname=restart queue=medium ctime=1565048821 qtime=1565048821 etime=1565048821 start=1565048829 exec\_host=compute-0-13/0_40+compute-0-14/0_40+compute-0-15/0_40+compute-0-16/0_40+compute-0-17/0_40+compute-0-8/0_40+compute-0-18/0_40+compute-0-19/0_40+compute-0-20/0_40+compute-0-21/0_40 exec\_vnode=(compute-0-13:ncpus=40)+(compute-0-14:ncpus=40)+(compute-0-15:ncpus=40)+(compute-0-16:ncpus=40)+(compute-0-17:ncpus=40)+(compute-0-8:ncpus=40)+(compute-0-18:ncpus=40)+(compute-0-19:ncpus=40)+(compute-0-20:ncpus=40)+(compute-0-21:ncpus=40) Resource\_List.mpiprocs=400 Resource\_List.ncpus=400 Resource\_List.nodect=10 Resource\_List.place=free Resource\_List.select=10:ncpus=40:mpiprocs=40 Resource\_List.walltime=01:00:00 session=0 end=1565048829 run\_count=20  
08/05/2019 16:47:09 A user=reinecke group=reinecke project=\_pbs\_project\_default jobname=restart queue=medium ctime=1565048821 qtime=1565048821 etime=1565048821 start=1565048829 exec\_host=compute-0-13/0_40+compute-0-14/0_40+compute-0-15/0_40+compute-0-16/0_40+compute-0-17/0_40+compute-0-8/0_40+compute-0-18/0_40+compute-0-19/0_40+compute-0-20/0_40+compute-0-21/0_40 exec\_vnode=(compute-0-13:ncpus=40)+(compute-0-14:ncpus=40)+(compute-0-15:ncpus=40)+(compute-0-16:ncpus=40)+(compute-0-17:ncpus=40)+(compute-0-8:ncpus=40)+(compute-0-18:ncpus=40)+(compute-0-19:ncpus=40)+(compute-0-20:ncpus=40)+(compute-0-21:ncpus=40) Resource\_List.mpiprocs=400 Resource\_List.ncpus=400 Resource\_List.nodect=10 Resource\_List.place=free Resource\_List.select=10:ncpus=40:mpiprocs=40 Resource\_List.walltime=01:00:00 resource\_assigned.ncpus=400  
08/05/2019 16:47:10 S Obit received momhop:21 serverhop:21 state:4 substate:41  
08/05/2019 16:47:10 S Discard running job, A sister Mom failed to delete job  
08/05/2019 16:47:10 S Job requeued, execution node down  
08/05/2019 16:47:10 A user=reinecke group=reinecke project=\_pbs\_project\_default jobname=restart queue=medium ctime=1565048821 qtime=1565048821 etime=1565048821 start=1565048829 exec\_host=compute-0-13/0_40+compute-0-14/0_40+compute-0-15/0_40+compute-0-16/0_40+compute-0-17/0_40+compute-0-8/0_40+compute-0-18/0_40+compute-0-19/0_40+compute-0-20/0_40+compute-0-21/0_40 exec\_vnode=(compute-0-13:ncpus=40)+(compute-0-14:ncpus=40)+(compute-0-15:ncpus=40)+(compute-0-16:ncpus=40)+(compute-0-17:ncpus=40)+(compute-0-8:ncpus=40)+(compute-0-18:ncpus=40)+(compute-0-19:ncpus=40)+(compute-0-20:ncpus=40)+(compute-0-21:ncpus=40) Resource\_List.mpiprocs=400 Resource\_List.ncpus=400 Resource\_List.nodect=10 Resource\_List.place=free Resource\_List.select=10:ncpus=40:mpiprocs=40 Resource\_List.walltime=01:00:00 session=0 end=1565048830 Exit\_status=-3 resources\_used.cpupercent=0 resources\_used.cput=00:00:00 resources\_used.mem=0kb resources\_used.ncpus=400 resources\_used.vmem=0kb resources\_used.walltime=00:00:00 run\_count=21  
08/05/2019 16:47:10 A user=reinecke group=reinecke project=\_pbs\_project\_default jobname=restart queue=medium ctime=1565048821 qtime=1565048821 etime=1565048821 start=1565048829 exec\_host=compute-0-13/0_40+compute-0-14/0_40+compute-0-15/0_40+compute-0-16/0_40+compute-0-17/0_40+compute-0-8/0_40+compute-0-18/0_40+compute-0-19/0_40+compute-0-20/0_40+compute-0-21/0_40 exec\_vnode=(compute-0-13:ncpus=40)+(compute-0-14:ncpus=40)+(compute-0-15:ncpus=40)+(compute-0-16:ncpus=40)+(compute-0-17:ncpus=40)+(compute-0-8:ncpus=40)+(compute-0-18:ncpus=40)+(compute-0-19:ncpus=40)+(compute-0-20:ncpus=40)+(compute-0-21:ncpus=40) Resource\_List.mpiprocs=400 Resource\_List.ncpus=400 Resource\_List.nodect=10 Resource\_List.place=free Resource\_List.select=10:ncpus=40:mpiprocs=40 Resource\_List.walltime=01:00:00 session=0 end=1565048830 run\_count=21  
08/05/2019 17:54:51 L Received qrun request  
08/05/2019 17:54:51 L Considering job to run  
08/05/2019 17:54:51 S Job Run at request of Scheduler@smaster1. on exec\_vnode (compute-0-0:ncpus=40)+(compute-0-1:ncpus=40)+(compute-0-2:ncpus=40)+(compute-0-3:ncpus=40)+(compute-0-4:ncpus=40)+(compute-0-5:ncpus=40)+(compute-0-6:ncpus=40)+(compute-0-7:ncpus=40)+(compute-0-9:ncpus=40)+(compute-0-10:ncpus=40)  
08/05/2019 17:54:51 S Job Modified at request of Scheduler@smaster1.  
08/05/2019 17:54:51 L Job run  
08/05/2019 17:54:51 A user=reinecke group=reinecke project=\_pbs\_project\_default jobname=restart queue=medium ctime=1565048821 qtime=1565048821 etime=0 start=1565052891 exec\_host=compute-0-0/0_40+compute-0-1/0_40+compute-0-2/0_40+compute-0-3/0_40+compute-0-4/0_40+compute-0-5/0_40+compute-0-6/0_40+compute-0-7/0_40+compute-0-9/0_40+compute-0-10/0_40 exec\_vnode=(compute-0-0:ncpus=40)+(compute-0-1:ncpus=40)+(compute-0-2:ncpus=40)+(compute-0-3:ncpus=40)+(compute-0-4:ncpus=40)+(compute-0-5:ncpus=40)+(compute-0-6:ncpus=40)+(compute-0-7:ncpus=40)+(compute-0-9:ncpus=40)+(compute-0-10:ncpus=40) Resource\_List.mpiprocs=400 Resource\_List.ncpus=400 Resource\_List.nodect=10 Resource\_List.place=free Resource\_List.select=10:ncpus=40:mpiprocs=40 Resource\_List.walltime=01:00:00 resource\_assigned.ncpus=400  
08/05/2019 18:55:28 S Obit received momhop:22 serverhop:22 state:4 substate:42  
08/05/2019 18:55:28 S Exit\_status=271 resources\_used.cpupercent=4008 resources\_used.cput=40:11:03 resources\_used.mem=4724692kb resources\_used.ncpus=400 resources\_used.vmem=38967208kb resources\_used.walltime=01:00:37  
08/05/2019 18:55:28 A user=reinecke group=reinecke project=\_pbs\_project\_default jobname=restart queue=medium ctime=1565048821 qtime=1565048821 etime=0 start=1565052891 exec\_host=compute-0-0/0_40+compute-0-1/0_40+compute-0-2/0_40+compute-0-3/0_40+compute-0-4/0_40+compute-0-5/0_40+compute-0-6/0_40+compute-0-7/0_40+compute-0-9/0_40+compute-0-10/0_40 exec\_vnode=(compute-0-0:ncpus=40)+(compute-0-1:ncpus=40)+(compute-0-2:ncpus=40)+(compute-0-3:ncpus=40)+(compute-0-4:ncpus=40)+(compute-0-5:ncpus=40)+(compute-0-6:ncpus=40)+(compute-0-7:ncpus=40)+(compute-0-9:ncpus=40)+(compute-0-10:ncpus=40) Resource\_List.mpiprocs=400 Resource\_List.ncpus=400 Resource\_List.nodect=10 Resource\_List.place=free Resource\_List.select=10:ncpus=40:mpiprocs=40 Resource\_List.walltime=01:00:00 session=178609 end=1565056528 Exit\_status=271 resources\_used.cpupercent=4008 resources\_used.cput=40:11:03 resources\_used.mem=4724692kb resources\_used.ncpus=400 resources\_used.vmem=38967208kb resources\_used.walltime=01:00:37 run\_count=22

---

<div class="post-metadata">

**Author:** ![brownwrap](https://avatars.discourse-cdn.com/v4/letter/b/94ad74/32.png) [@brownwrap](https://community.openpbs.org/u/brownwrap)\
**Post date:** [August 6, 2019, 10:19pm UTC](https://community.openpbs.org/t/execution-node-down/1723/6 "2019-08-06T22:19:45Z")

</div>

’

The system is not letting upload regular text files.

---

<div class="post-metadata">

**Author:** ![brownwrap](https://avatars.discourse-cdn.com/v4/letter/b/94ad74/32.png) [@brownwrap](https://community.openpbs.org/u/brownwrap)\
**Post date:** [August 7, 2019, 7:04pm UTC](https://community.openpbs.org/t/execution-node-down/1723/7 "2019-08-07T19:04:28Z")

</div>

Did the output of tracejob provides and additional info? I have made note of the suggested way to offline a node, but what is the difference between the method you provides and:

pbsnode -o node  
pbsnode -c node

Thanks.

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [August 9, 2019, 7:48am UTC](https://community.openpbs.org/t/execution-node-down/1723/8 "2019-08-09T07:48:37Z")

</div>

Both are the same, but doing it via qmgr way is less error prone. I will update this thread once i get to know the reason. Probably, the node incarnation is updated correctly in the datastore and memory when we use qmgr way. Trust rest is working for you, thanks for sharing the above logs.

08/05/2019 16:47:10 S Discard running job, A sister Mom failed to delete job  
08/05/2019 16:47:10 S Job requeued, execution node down

Please make sure when you bring up the compute nodes, the jobs spool directory are clean.
