Before fail_fs, cephadm had to reduce max_mds to 1 to avoid having two active MDS modifying on-disk structures with new versions, communicating cross-version-incompatible messages, or other potential incompatibilities. Scaling down that way requires graceful rank deactivation and migration, which under heavy load can take far too long, stall, or severely impact clients.
fail_fs lets cephadm keep max_mds > 1 and instead issue fs fail: the monitor marks the filesystem not-joinable and forcibly fails all ranks at once, blocklisting each MDS, so the daemons drop straight to standby without the slow graceful stop. cephadm also bypasses the per-MDS ok-to-stop gate in this mode, removing the single-rank serialization. The daemons are still redeployed in batches rather than literally all at once, but the upgrade window is shorter, the period of unavailability is reduced, and the process is less disruptive overall.
And at the end cephadm just sets the filesystem joinable again rather than having to scale max_mds back up.
Cheers,
Frédéric.
--
Frédéric Nass
Ceph Ambassador France | Senior Ceph Engineer @ CLYSO
Check our online Config Diff Tool, it's great!