# Pbs doesnt start after openhpc update

**URL:** <https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989>\
**Category:** Users/Site Administrators\
**Created:** [May 10, 2018, 8:38pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989 "2018-05-10T20:38:46Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![trumee](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@trumee](https://community.openpbs.org/u/trumee)\
**Post date:** [May 10, 2018, 8:38pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/1 "2018-05-10T20:38:46Z")

</div>

Hello,

I updated the OpenHPC server and unfortunately PBS is refusing to start.

```
# systemctl status pbs
● pbs.service - Portable Batch System
   Loaded: loaded (/opt/pbs/libexec/pbs_init.d; enabled; vendor preset: disabled)
   Active: failed (Result: exit-code) since Thu 2018-05-10 15:24:56 CDT; 12min ago
     Docs: man:pbs(8)
  Process: 2811 ExecStart=/opt/pbs/libexec/pbs_init.d start (code=exited, status=1/FAILURE)

May 10 15:24:53 hpc.localdomain pbs_init.d[2811]: Running /opt/pbs/libexec/pbs_habitat to update it.
May 10 15:24:53 hpc.localdomain pbs_init.d[2811]: ***
May 10 15:24:53 hpc.localdomain su[3087]: (to postgres) root on none
May 10 15:24:56 hpc.localdomain su[5209]: (to postgres) root on none
May 10 15:24:56 hpc.localdomain pbs_init.d[2811]: /opt/pbs/pgsql/pg_upgrade not found
May 10 15:24:56 hpc.localdomain pbs_init.d[2811]: Failed to upgrade PBS Datastore
May 10 15:24:56 hpc.localdomain systemd[1]: pbs.service: control process exited, code=exited status=1
May 10 15:24:56 hpc.localdomain systemd[1]: Failed to start Portable Batch System.
May 10 15:24:56 hpc.localdomain systemd[1]: Unit pbs.service entered failed state.
May 10 15:24:56 hpc.localdomain systemd[1]: pbs.service failed.

```

Any idea what could be the issue?

---

<div class="post-metadata">

**Author:** ![trumee](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@trumee](https://community.openpbs.org/u/trumee)\
**Post date:** [May 10, 2018, 8:51pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/2 "2018-05-10T20:51:43Z")

</div>

I installed postgresql-upgrade-9.2.23-3.el7\_4.x86\_64 and now have,

```
# which pg_upgrade
/usr/bin/pg_upgrade

# rpm -qa postgresql*
postgresql-server-9.2.23-3.el7_4.x86_64
postgresql-jdbc-9.2.1002-5.el7.noarch
postgresql-9.2.23-3.el7_4.x86_64
postgresql-upgrade-9.2.23-3.el7_4.x86_64
postgresql-libs-9.2.23-3.el7_4.x86_64

```

I dont see a /opt/pbs/pgsql directory as requested by the service (pbs\_init.d[22921]: /opt/pbs/pgsql/pg\_upgrade not found),

```
ls -la /opt/pbs/
total 40
drwxr-xr-x 10 root root 4096 Mar 5 12:21 .
drwxr-xr-x 7 root root 4096 Apr 10 23:59 ..
drwxr-xr-x 2 root root 4096 May 10 13:15 bin
drwxr-xr-x 2 root root 4096 May 10 13:15 etc
drwxr-xr-x 2 root root 4096 May 10 13:15 include
drwxr-xr-x 5 root root 4096 May 10 13:15 lib
drwxr-xr-x 2 root root 4096 May 10 13:15 libexec
drwxr-xr-x 2 root root 4096 May 10 13:15 sbin
drwxr-xr-x 3 root root 4096 Mar 5 12:21 share
drwxr-xr-x 3 root root 4096 May 10 13:15 unsupported

# rpm -qa pbspro*
pbspro-server-ohpc-14.1.2-9.1.x86_64

```

How do i get the /opt/pbs/pgsql/ directory?

---

<div class="post-metadata">

**Author:** ![trumee](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@trumee](https://community.openpbs.org/u/trumee)\
**Post date:** [May 10, 2018, 9:14pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/3 "2018-05-10T21:14:15Z")

</div>

I found this [thread](http://community.openpbs.org/t/issue-starting-pbs/403) related to my problem. The suggestion there was to delete /var/spool/pbs and start afresh. Is it possible to avoid this, i would like to get the information on the jobs run for the past week for accounting purposes.

---

<div class="post-metadata">

**Author:** ![trumee](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@trumee](https://community.openpbs.org/u/trumee)\
**Post date:** [May 10, 2018, 11:00pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/4 "2018-05-10T23:00:47Z")

</div>

The function responsible to update the db is in /opt/pbs/libexec/pbs\_habitat

```
upgrade_db() {
        if [! -x "${inst_dir}/bin/pg_upgrade"]; then
                echo "${inst_dir}/pg_upgrade not found"
                return 1
        fi

```

Why is this using inst\_dir which is set to /opt/pbs instead of the Postgres binary path? The pgsql environments is set from /opt/pbs/libexec/pbs\_pgsql\_env.sh, which sets the PGSQL\_DIR variable correctly.

Is the upgrade not been updated?

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [May 11, 2018, 2:28pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/5 "2018-05-11T14:28:05Z")

</div>

> [@trumee](#):
>
> The suggestion there was to delete /var/spool/pbs and start afresh. Is it possible to avoid this, i would like to get the information on the jobs run for the past week for accounting purposes.

You could backup the /var/spool/pbs/server\_priv/accounting folder, this has all the accounting information of the jobs that are run on the cluster till the PBS was active.

Once you have backed up , you can delete and re-initiate it again. Copy back the accounting logs back to the same location , if required.

---

<div class="post-metadata">

**Author:** ![trumee](https://avatars.discourse-cdn.com/v4/letter/t/848f3c/32.png) [@trumee](https://community.openpbs.org/u/trumee)\
**Post date:** [May 11, 2018, 2:37pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/6 "2018-05-11T14:37:34Z")

</div>

I downgraded to pbs-pro-server-ohpc-14.1.0-30.2.x86\_64 from 14.1.2-9.1.x86\_64, and have working PBS now. What directories do i need to delete in /var/spool/pbs to re-initialize after upgrading to 14.1.2-9.1 ?

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [May 11, 2018, 2:40pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/7 "2018-05-11T14:40:27Z")

</div>

Please follow these steps from @mkaro

> [@Can not delete a vnode whose name prefix with #(hash)](http://community.openpbs.org/t/can-not-delete-a-vnode-whose-name-prefix-with-hash/939/4):
>
> You may use /opt/pbs/sbin/pbs\_ds\_passwd to change the password for the psql database. The remaining steps that @adarsh outlined are correct. If you don’t mind starting from scratch, you may also do the following… Stop all PBS Pro services Delete /var/spool/pbs (the PBS\_HOME directory) Start PBS and reconfigure the services Keep in mind, this process will delete all configuration data from PBS Pro and you’ll effectively be starting with a fresh installation. Use this as a last resort.

---

<div class="post-metadata">

**Author:** ![mkaro](https://yyz2.discourse-cdn.com/flex030/user_avatar/community.openpbs.org/mkaro/32/85_2.png) [@mkaro](https://community.openpbs.org/u/mkaro)\
**Post date:** [May 11, 2018, 5:25pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/8 "2018-05-11T17:25:16Z")

</div>

Hello @trumee,

It appears there has already been a ticket filed for this… [https://pbspro.atlassian.net/browse/PP-756](https://pbspro.atlassian.net/browse/PP-756)

The bug has been escalated and we will address it ASAP. Thank you for pointing it out.

Mike

---

<div class="post-metadata">

**Author:** ![smilezy](https://avatars.discourse-cdn.com/v4/letter/s/57b2e6/32.png) [@smilezy](https://community.openpbs.org/u/smilezy)\
**Post date:** [November 1, 2018, 1:54am UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/9 "2018-11-01T01:54:17Z")

</div>

Hello adarsh,  
“You could backup the /var/spool/pbs/server\_priv/accounting folder, this has all the accounting information of the jobs that are run on the cluster till the PBS was active.”，but I can’t find the value of job’s Output\_Path attribute，can find it using qstat -f.From which file we can find the job’s Output\_Path property value for job?

---

<div class="post-metadata">

**Author:** ![adarsh](https://avatars.discourse-cdn.com/v4/letter/a/f07891/32.png) [@adarsh](https://community.openpbs.org/u/adarsh)\
**Post date:** [November 1, 2018, 12:53pm UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/10 "2018-11-01T12:53:22Z")

</div>

- It is in the Job History ( Memory ) - depends on the job\_history\_enable and job\_history\_duration attributes in the server configuration.
- It is in the PBS datastore - also depends on the job\_history\_enable and job\_history\_duration attributes in the server configuration.

The Output\_Path is not found in the accounting logs ☹

qstat -fx # should work for the job in the history and should get the Output\_Path

---

<div class="post-metadata">

**Author:** ![smilezy](https://avatars.discourse-cdn.com/v4/letter/s/57b2e6/32.png) [@smilezy](https://community.openpbs.org/u/smilezy)\
**Post date:** [November 2, 2018, 1:11am UTC](https://community.openpbs.org/t/pbs-doesnt-start-after-openhpc-update/989/11 "2018-11-02T01:11:34Z")

</div>

Thanks for your patient reply.
