A brief period of service unavailability would be tolerable for our environment.
Our primary concern is the behavior of the clients during and after the failover/upgrade process. Specifically, we would like to understand whether kernel CephFS clients can reliably recover once the filesystem becomes available again, or whether there are known scenarios in which client nodes may remain stuck, encounter kernel hangs, or require a reboot to recover.
We are currently using the Linux kernel CephFS client.
At 2026-06-01 22:33:09, "Anthony D'Atri via dev" <dev@ceph.io> wrote:
Merci.Do you have any sense of how long the unavailability will be for a volume with, say 8 large MDS instances?During such an operation, do the clients wedge and come back, or do they error?This looks to have been introduced in Reef. Since the Manager is updated before MDSes, would one reasonably expect to be able to use this feature during an upgrade *to* Reef?I would like to add context to the eocs.On Jun 1, 2026, at 4:35 AM, Frédéric Nass via dev <dev@ceph.io> wrote:Hi Anthony,The explanation is here: https://tracker.ceph.com/issues/55715 (PR https://github.com/ceph/ceph/pull/47756).
Before fail_fs, cephadm had to reduce max_mds to 1 to avoid having two active MDS modifying on-disk structures with new versions, communicating cross-version-incompatible messages, or other potential incompatibilities. Scaling down that way requires graceful rank deactivation and migration, which under heavy load can take far too long, stall, or severely impact clients.
fail_fs lets cephadm keep max_mds > 1 and instead issue fs fail: the monitor marks the filesystem not-joinable and forcibly fails all ranks at once, blocklisting each MDS, so the daemons drop straight to standby without the slow graceful stop. cephadm also bypasses the per-MDS ok-to-stop gate in this mode, removing the single-rank serialization. The daemons are still redeployed in batches rather than literally all at once, but the upgrade window is shorter, the period of unavailability is reduced, and the process is less disruptive overall.And at the end cephadm just sets the filesystem joinable again rather than having to scale max_mds back up.Cheers,Frédéric.--
Frédéric NassCeph Ambassador France | Senior Ceph Engineer @ CLYSOCheck our online Config Diff Tool, it's great!Le ven. 29 mai 2026 à 17:01, Anthony D'Atri via dev <dev@ceph.io> a écrit :I’ve struggled a bit understanding this strategy. Help me understand please how completely disabling the filesystem / volume is more tolerable than reducing to one?On May 29, 2026, at 4:30 AM, Dhairya Parmar via dev <dev@ceph.io> wrote:Hi,There is a solution for this (which works by failing the filesystem), please check the note here - https://docs.ceph.com/en/latest/cephadm/upgrade/#starting-the-upgradeDhairya ParmarSoftware Engineer, CephFSOn Fri, May 29, 2026 at 3:17 PM 刘先生 via dev <dev@ceph.io> wrote:hi:Our Ceph cluster is deployed using Ceph cephadm and uses multiple MDS daemons. The documentation says that during a cephadm upgrade, the filesystem should be reduced to a single active MDS. However, this is not feasible for us(single mds have not enough mem). Is there a more elegant way to perform the upgrade?
Thank you very much.