Hi,
I was wondering (and could not find a solution yet) how to properly and gracefully shut down an Incus cluster? This is relevant both for planned maintenance events where the cluster shall be stopped manually, but also to plan and implement power-down logic in e.g. UPS scripts, to prevent an unclean cluster shutdown in case of power issues.
I thought about a few aspects:
- Evacuating cluster members does not help because it will trigger unwanted/unnecessary instance migrations.
- Stopping all instances on a specific member is surprisingly tedious, esp. when using projects: There is no
incus stop --all-projects --only-local-onesor something similar, so you need to filter for running instances on a specific member and iterate over all projects. - Even then, after stopping instances manually and issuing
incus admin shutdownmember by member, I noticed my cluster became unresponsive when the quorum was lost (in my case: only 1 member remaining of 3). So very unclear if dqlite data is lost.
To expand the question, I am also wondering how to combine an Incus shutdown with a proper Ceph shutdown, considering it’s an all-in-one HCI cluster. There are best practices like ceph osd set noout (and nobackfill, norecover), but to start the cluster again, you would need to unset those before Incus starts, how?
Any thoughts or input is appreciated.
Regards