Hi Frédérik,
could my thread
(https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/message/CSVJC2WLQ3V4RTWQF5GU4JAPEA5GJDHJ/) be related to this as well? I don't think I tested it with MDS daemons, but from my perspective there's definitely something wrong with staggered
upgrades.
I don't mean to hijack this thread, just wanted to bring some
attention to the topic of staggered upgrades.
Thanks
Eugen
Zitat von Frédéric Nass via dev <dev@ceph.io>:
> Hey folks,
>
> I've been testing the fail_fs option over the last few hours. I discovered
> that during an MDS upgrade — even when restricting the scope to a single
> filesystem with ceph orch upgrade ... --services or --daemon-types — all
> filesystems are either failed (if fail_fs is set to true) or dropped to
> max_mds 1. This is highly disruptive and rules out any smooth, worry-free
> upgrade scenario for clusters with multiple filesystems.
>
> I've opened a tracker for this: https://tracker.ceph.com/issues/77074.
>
> --
> Frédéric Nass
> Ceph Ambassador France | Senior Ceph Engineer @ CLYSO
> Try our Ceph Analyzer -- https://analyzer.clyso.com/
> https://clyso.com | frederic.nass@clyso.com
>
> Le mar. 2 juin 2026 à 14:16, Dhairya Parmar via dev <dev@ceph.io> a écrit :
>
>> Hi folks,
>> (replies in-line)
>>
>>
>> On Mon, Jun 1, 2026 at 7:49 PM Anthony D'Atri <aad@dreamsnake.net> wrote:
>>
>>> Merci.
>>>
>>> Do you have any sense of how long the unavailability will be for a volume
>>> with, say 8 large MDS instances?
>>>
>>>
>> Cephadm upgrades daemons one-by-one, which will indeed increase the time
>> linearly. However, the non-trivial aspect, from the MDS point of view, is
>> how quickly the MDS transitions from `up:replay` -> `up:resolve` ->
>> `up:reconnect` -> `up:rejoin` -> `up:active`. If you start from the first
>> phase (journal replay) - it is this part that usually consumes a big chink
>> of time because it depends on how big the journal is unflushed (the busier
>> the MDS, the more chance of having a bigger journal == more time spent
>> during replay). The other states that should contribute to delays are
>> `up:resolve` (since it is an active-active MDS setup) and `up:rejoin`.
>> Source: https://docs.ceph.com/en/latest/cephfs/mds-states
>>
>>
>>> During such an operation, do the clients wedge and come back, or do they
>>> error?
>>>
>>
>> Given the design, there is no client-side force-session-close timeout for
>> FUSE/kclient, so, clients would just stall. This depends on the path: if it
>> is the data path and the client has read/write caps, I/O with OSDs would
>> continue without trouble. However, for any metadata exchange or anything
>> that requires comms with MDS -- such as a cap change/upgrade/flush, opening
>> a new file or flushing the caps (fsync or close) and so on -- I think the
>> client will likely hang or stall. For this particular cluster which uses
>> only kclient, I think it would transition to the D state (uninterruptible
>> sleep) until an active MDS responds.
>>
>>
>>>
>>> This looks to have been introduced in Reef. Since the Manager is updated
>>> before MDSes, would one reasonably expect to be able to use this feature
>>> during an upgrade *to* Reef?
>>>
>>
>> During a staggered upgrade, yes. The mgr needs to be upgraded to Reef
>> first and then the MDSes should be upgraded.
>>
>>
>>>
>>> I would like to add context to the eocs.
>>>
>>>
>>> On Jun 1, 2026, at 4:35 AM, Frédéric Nass via dev <dev@ceph.io> wrote:
>>>
>>> Hi Anthony,
>>>
>>> The explanation is here: https://tracker.ceph.com/issues/55715 (PR
>>> https://github.com/ceph/ceph/pull/47756).
>>>
>>> Before fail_fs, cephadm had to reduce max_mds to 1 to avoid having two
>>> active MDS modifying on-disk structures with new versions, communicating
>>> cross-version-incompatible messages, or other potential incompatibilities.
>>> Scaling down that way requires graceful rank deactivation and migration,
>>> which under heavy load can take far too long, stall, or severely impact
>>> clients.
>>>
>>> fail_fs lets cephadm keep max_mds > 1 and instead issue fs fail: the
>>> monitor marks the filesystem not-joinable and forcibly fails all ranks at
>>> once, blocklisting each MDS, so the daemons drop straight to standby
>>> without the slow graceful stop. cephadm also bypasses the per-MDS
>>> ok-to-stop gate in this mode, removing the single-rank serialization. The
>>> daemons are still redeployed in batches rather than literally all at once,
>>> but the upgrade window is shorter, the period of unavailability is reduced,
>>> and the process is less disruptive overall.
>>>
>>> And at the end cephadm just sets the filesystem joinable again rather
>>> than having to scale max_mds back up.
>>>
>>> Cheers,
>>> Frédéric.
>>>
>>> --
>>> Frédéric Nass
>>> Ceph Ambassador France | Senior Ceph Engineer @ CLYSO
>>> Check our online Config Diff Tool, it's great!
>>> https://analyzer.clyso.com/#/analyzer/config-diff
>>> https://clyso.com | frederic.nass@clyso.com
>>>
>>>
>>> Le ven. 29 mai 2026 à 17:01, Anthony D'Atri via dev <dev@ceph.io> a
>>> écrit :
>>>
>>>> I’ve struggled a bit understanding this strategy. Help me understand
>>>> please how completely disabling the filesystem / volume is more tolerable
>>>> than reducing to one?
>>>>
>>>> On May 29, 2026, at 4:30 AM, Dhairya Parmar via dev <dev@ceph.io> wrote:
>>>>
>>>>
>>>> Hi,
>>>> There is a solution for this (which works by failing the filesystem),
>>>> please check the note here -
>>>> https://docs.ceph.com/en/latest/cephadm/upgrade/#starting-the-upgrade
>>>>
>>>> *Dhairya Parmar*
>>>> Software Engineer, CephFS
>>>>
>>>>
>>>> On Fri, May 29, 2026 at 3:17 PM 刘先生 via dev <dev@ceph.io> wrote:
>>>>
>>>>> hi:
>>>>> Our Ceph cluster is deployed using Ceph cephadm and uses multiple MDS
>>>>> daemons. The documentation says that during a cephadm upgrade, the
>>>>> filesystem should be reduced to a single active MDS. However, this is not
>>>>> feasible for us(single mds have not enough mem). Is there a more elegant
>>>>> way to perform the upgrade?
>>>>> Thank you very much.
>>>>>
>>>>>
>>>