Gracefully shutting down a cluster for maintenance

Hi,

I was wondering (and could not find a solution yet) how to properly and gracefully shut down an Incus cluster? This is relevant both for planned maintenance events where the cluster shall be stopped manually, but also to plan and implement power-down logic in e.g. UPS scripts, to prevent an unclean cluster shutdown in case of power issues.

I thought about a few aspects:

  • Evacuating cluster members does not help because it will trigger unwanted/unnecessary instance migrations.
  • Stopping all instances on a specific member is surprisingly tedious, esp. when using projects: There is no incus stop --all-projects --only-local-ones or something similar, so you need to filter for running instances on a specific member and iterate over all projects.
  • Even then, after stopping instances manually and issuing incus admin shutdown member by member, I noticed my cluster became unresponsive when the quorum was lost (in my case: only 1 member remaining of 3). So very unclear if dqlite data is lost.

To expand the question, I am also wondering how to combine an Incus shutdown with a proper Ceph shutdown, considering it’s an all-in-one HCI cluster. There are best practices like ceph osd set noout (and nobackfill, norecover), but to start the cluster again, you would need to unset those before Incus starts, how?

Any thoughts or input is appreciated.

Regards

incus cluster evacuate --action=stop

Incus will eventually timeout on shutdown when without a quorum, so once you have all your servers in evacuated state, you can shut them down, all of them except for the last two will shutdown quickly, the remaining two will unblock after a while of trying to get back to quorum.

Thank you, Stéphane, hadn’t had the evacuate --action=stop on my radar! I will run some tests.

Regards

It would be nice to allow a quicker shutdown of the last members. I may have encountered situations where main power loss required a clean full shutdown in just a few minutes, and having those last servers taking their time isn’t great.

Good point, waiting on arbitrary timeouts is not feasible in a power outage scenario where you’d want controlled shutdown.

Would be interesting to see what that looks like once all instances are stopped.
With recent tweaks to the database layer, I’m not seeing much more than a 30s or so delay in there in theory.