Recover of ovn network

Okay, I don’t recall how much we allow users to mess with this stuff, so it will probably fail, but try:

  • incus config trust remove 74ff
  • incus config trust add-certificate --type=server --name=tnode4 /path/to/tnode4/server.crt

That needs to be done from one of the working servers.

Thanks for the info but getting this error message.
indiana@tnode1:~$ incus config trust add-certificate --type=server --name=tnode4 /var/lib/incus/server.crt
Error: Unknown certificate type β€œserver”

okay, add it with type=client and then change it in the DB after the fact :slight_smile:
That will need a restart of Incus on the servers to pick up the new type though.

incus admin sql global "UPDATE certificates SET type=2 WHERE name='tnode4'" once you have it added as a client certificate

Here are the outputs:

indiana@tnode1:~$ incus config trust remove 74ff
Error: failed to notify peer 192.168.1.204:8443: Daemon is starting up
indiana@tnode1:~$ incus config trust remove 74ff
Error: Certificate not found

indiana@tnode1:~$ incus config trust list
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚      NAME      β”‚  TYPE  β”‚ DESCRIPTION β”‚ FINGERPRINT  β”‚     EXPIRY DATE      β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ indiana@debian β”‚ client β”‚             β”‚ 22a880ef01b1 β”‚ 2036/01/11 20:45 +03 β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ tnode1         β”‚ server β”‚             β”‚ 08488766685b β”‚ 2036/05/18 08:35 +03 β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ tnode2         β”‚ server β”‚             β”‚ 7d7e2c3c01fc β”‚ 2036/05/18 09:52 +03 β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ tnode3         β”‚ server β”‚             β”‚ 98787c2fb56f β”‚ 2036/05/18 11:23 +03 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
indiana@tnode1:~$ incus config trust add-certificate --type=client --name=tnode4 /var/lib/incus/server.crt
Error: Certificate already in trust store
indiana@tnode1:~$ incus admin sql global "UPDATE certificates SET type=2 WHERE name='tnode4'"
Rows affected: 0
indiana@tnode1:~$

You need to add-certificate tnode4’s /var/lib/incus/server.crt, above you’re trying to re-insert tnode1’s certificate instead.

My bad, I have added the new certificate, seems that the previous alarms are gone but one last struggle. How can I add/remove the β€˜local/backuppool’ using sql?.

Regards.

Sep 09 22:24:31 tnode4 systemd[1]: Starting incus.service - Incus - Daemon...
Sep 09 22:24:32 tnode4 incusd[1772]: time="2026-09-09T22:24:32+03:00" level=warning msg="QMP monitor read failed" err="read unix @->/tmp/599661677: use of closed network connection"
Sep 09 22:25:06 tnode4 incusd[1772]: time="2026-09-09T22:25:06+03:00" level=error msg="Failed to start the daemon" err="Failed to mount backups storage: Failed to mount storage volume "local/backuppool": Storage volume "backuppool" in project "default" of typ>
Sep 09 22:25:07 tnode4 incusd[1772]: Error: Failed to mount backups storage: Failed to mount storage volume "local/backuppool": Storage volume "backuppool" in project "default" of type "custom" does not exist on pool "local": Storage volume not found
Sep 09 22:25:07 tnode4 systemd[1]: incus.service: Main process exited, code=exited, status=1/FAILURE

Hmm, your cluster is in a really weird state…

Can you do sqlite3 /var/lib/incus/database/local.db "SELECT * FROM config"?

root@tnode4:/var/log/incus# sqlite3 /var/lib/incus/database/local.db "SELECT * FROM config"
1|cluster.https_address|192.168.1.204:8443
2|core.https_address|192.168.1.204:8443
3|storage.backups_volume|local/backuppool
4|storage.images_volume|local/imagepool

Please also the sqlite3 command on some of the healthy nodes of the cluster.

Not just from the β€œbroken” tnode4.

Here are the other nodes:

indiana@tnode1:~$ sudo sqlite3 /var/lib/incus/database/local.db "SELECT * FROM config"
[sudo] password for indiana:
1|cluster.https_address|192.168.1.201:8443
2|core.https_address|192.168.1.201:8443
3|core.metrics_address|:8444
4|storage.backups_volume|local/backuppool
5|storage.images_volume|local/imagepool

indiana@tnode3:~$ sudo sqlite3 /var/lib/incus/database/local.db "SELECT * FROM config"
[sudo] password for indiana:
1|core.https_address|192.168.1.203:8443
2|cluster.https_address|192.168.1.203:8443
3|storage.backups_volume|local/backuppool
4|storage.images_volume|local/imagepool

indiana@tnode2:~$ sudo sqlite3 /var/lib/incus/database/local.db "SELECT * FROM config"
[sudo] password for indiana:
1|cluster.https_address|192.168.1.202:8443
2|core.https_address|192.168.1.202:8443
5|storage.backups_volume|local/backuppool
6|storage.images_volume|local/imagepool

Given the database has apparently no record of those volumes for tnode4, easiest is to clear them for now.

sqlite3 /var/lib/incus/database/local.db "DELETE FROM config WHERE key IN ('storage.backups_volume', 'storage.images_volume')" should do the trick and allow the daemon to start again (run that only on tnode4).

That will likely leave actual volumes behind though but we’ll be able to see just how messed up the global database is at that point… I’m a bit fearful of records of tnode4 having somehow been cleared from the global database which would explain those missing volumes but could then also extend to missing instances and such.

Thank you very much for baring with me and helping me. Unlink those backuppool and imagepool then the node recovered.

Regards.