Re: how to upgrade multi mds but not change it to single mds?
Done: https://tracker.ceph.com/issues/77093 Zitat von Frédéric Nass via dev <dev@ceph.io>:
Hi Eugen,
I remember that thread, yes. The inconsistent behavior of --limit could come from elsewhere in the code, but likely there's something to investigate. Could you make a tracker for that?
Cheers, Frédéric.
Frédéric Nass
Senior Ceph Engineer
Ceph Ambassador, France
+49 89 215252-751 <https://call.ctrlq.org/+49%2089%20215252-751>
frederic.nass@clyso.com
www.clyso.com
Hohenzollernstr. 27, 80801 Munich
Utting a. A. | HR: Augsburg | HRB: 25866 | USt. ID-Nr.: DE2754306
Le mar. 2 juin 2026 à 16:13, Eugen Block via dev <dev@ceph.io> a écrit :
Hi Frédérik,
could my thread ( https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/message/CSVJC2WLQ3V...) be related to this as well? I don't think I tested it with MDS daemons, but from my perspective there's definitely something wrong with staggered upgrades.
I don't mean to hijack this thread, just wanted to bring some attention to the topic of staggered upgrades.
Thanks Eugen
Zitat von Frédéric Nass via dev <dev@ceph.io>:
Hey folks,
I've been testing the fail_fs option over the last few hours. I discovered that during an MDS upgrade — even when restricting the scope to a single filesystem with ceph orch upgrade ... --services or --daemon-types — all filesystems are either failed (if fail_fs is set to true) or dropped to max_mds 1. This is highly disruptive and rules out any smooth, worry-free upgrade scenario for clusters with multiple filesystems.
I've opened a tracker for this: https://tracker.ceph.com/issues/77074.
-- Frédéric Nass Ceph Ambassador France | Senior Ceph Engineer @ CLYSO Try our Ceph Analyzer -- https://analyzer.clyso.com/ https://clyso.com | frederic.nass@clyso.com
Le mar. 2 juin 2026 à 14:16, Dhairya Parmar via dev <dev@ceph.io> a écrit :
Hi folks, (replies in-line)
On Mon, Jun 1, 2026 at 7:49 PM Anthony D'Atri <aad@dreamsnake.net> wrote:
Merci.
Do you have any sense of how long the unavailability will be for a volume with, say 8 large MDS instances?
Cephadm upgrades daemons one-by-one, which will indeed increase the time linearly. However, the non-trivial aspect, from the MDS point of view, is how quickly the MDS transitions from `up:replay` -> `up:resolve` -> `up:reconnect` -> `up:rejoin` -> `up:active`. If you start from the first phase (journal replay) - it is this part that usually consumes a big chink of time because it depends on how big the journal is unflushed (the busier the MDS, the more chance of having a bigger journal == more time spent during replay). The other states that should contribute to delays are `up:resolve` (since it is an active-active MDS setup) and `up:rejoin`. Source: https://docs.ceph.com/en/latest/cephfs/mds-states
During such an operation, do the clients wedge and come back, or do they error?
Given the design, there is no client-side force-session-close timeout for FUSE/kclient, so, clients would just stall. This depends on the path: if it is the data path and the client has read/write caps, I/O with OSDs would continue without trouble. However, for any metadata exchange or anything that requires comms with MDS -- such as a cap change/upgrade/flush, opening a new file or flushing the caps (fsync or close) and so on -- I think the client will likely hang or stall. For this particular cluster which uses only kclient, I think it would transition to the D state (uninterruptible sleep) until an active MDS responds.
This looks to have been introduced in Reef. Since the Manager is
updated
before MDSes, would one reasonably expect to be able to use this feature during an upgrade *to* Reef?
During a staggered upgrade, yes. The mgr needs to be upgraded to Reef first and then the MDSes should be upgraded.
I would like to add context to the eocs.
On Jun 1, 2026, at 4:35 AM, Frédéric Nass via dev <dev@ceph.io> wrote:
Hi Anthony,
The explanation is here: https://tracker.ceph.com/issues/55715 (PR https://github.com/ceph/ceph/pull/47756).
Before fail_fs, cephadm had to reduce max_mds to 1 to avoid having two active MDS modifying on-disk structures with new versions,
communicating
cross-version-incompatible messages, or other potential incompatibilities. Scaling down that way requires graceful rank deactivation and migration, which under heavy load can take far too long, stall, or severely impact clients.
fail_fs lets cephadm keep max_mds > 1 and instead issue fs fail: the monitor marks the filesystem not-joinable and forcibly fails all ranks at once, blocklisting each MDS, so the daemons drop straight to standby without the slow graceful stop. cephadm also bypasses the per-MDS ok-to-stop gate in this mode, removing the single-rank serialization. The daemons are still redeployed in batches rather than literally all at once, but the upgrade window is shorter, the period of unavailability is reduced, and the process is less disruptive overall.
And at the end cephadm just sets the filesystem joinable again rather than having to scale max_mds back up.
Cheers, Frédéric.
-- Frédéric Nass Ceph Ambassador France | Senior Ceph Engineer @ CLYSO Check our online Config Diff Tool, it's great! https://analyzer.clyso.com/#/analyzer/config-diff https://clyso.com | frederic.nass@clyso.com
Le ven. 29 mai 2026 à 17:01, Anthony D'Atri via dev <dev@ceph.io> a écrit :
I’ve struggled a bit understanding this strategy. Help me understand please how completely disabling the filesystem / volume is more tolerable than reducing to one?
On May 29, 2026, at 4:30 AM, Dhairya Parmar via dev <dev@ceph.io> wrote:
Hi, There is a solution for this (which works by failing the filesystem), please check the note here - https://docs.ceph.com/en/latest/cephadm/upgrade/#starting-the-upgrade
*Dhairya Parmar* Software Engineer, CephFS
On Fri, May 29, 2026 at 3:17 PM 刘先生 via dev <dev@ceph.io> wrote:
> hi: > Our Ceph cluster is deployed using Ceph cephadm and uses multiple MDS > daemons. The documentation says that during a cephadm upgrade, the > filesystem should be reduced to a single active MDS. However, this is not > feasible for us(single mds have not enough mem). Is there a more elegant > way to perform the upgrade? > Thank you very much. > >
participants (1)
-
Eugen Block