Massive mon slow-ops when restarting mon daemons
Hey there! I’m currently in the process of upgrading the production cluster from v18.2.8 -> v19.2.6 via „ceph orch upgrade“. I tested this upgrade in our staging environment and everything went great. On our production environment I encounter some worrisome issues with the mon daemons. The cluster itself: 5x MON Daemons 5x MGR Daemons 1500x OSD - 2 larger Pools with each 16k in PG. … so a somewhat larger cluster. The orchestrator started re-deploying the MGR, which went fine. During the MON redeployment, the cluster went into warning with messages like „34846 slow ops. mon.mon1 has slow ops“. The counter climbed rapidly. Only way to stop this was to restart the MON in question. Sometimes I had to restart the MON multiple times. Afterwards the cluster is healthy again….until I restart any MON. The symptoms return until restarting fixes it again. I tried to dump the mon_ops_in_flight, but the command just hangs (probably because of the large error count). Filtering the journal for the issues: ``` journalctl -g "oldest is" --since "3 hours ago" -u ceph-d8a6f365-99de-42f2-b32b-c9ea11df945c@mon.partnach <mailto:ceph-d8a6f365-99de-42f2-b32b-c9ea11df945c@mon.partnach> Sep 14 09:45:05 partnach conmon[1735]: 2026-09-14T07:45:05.331+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 18286 slow ops, oldest is log(1000 entries from seq 2 at 2026-08-26T04:51:47.819315+0000) Sep 14 09:45:10 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 18505 slow ops, oldest is log(1000 entries from seq 4003 at 2026-09-11T12:25:52.136679+0000) Sep 14 09:45:10 partnach conmon[1735]: 2026-09-14T07:45:10.827+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 18505 slow ops, oldest is log(1000 entries from seq 4003 at 2026-09-11T12:25:52.136679+0000) Sep 14 09:45:15 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 18671 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-07T16:47:58.529125+0000) Sep 14 09:45:15 partnach conmon[1735]: 2026-09-14T07:45:15.924+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 18671 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-07T16:47:58.529125+0000) Sep 14 09:45:21 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 18814 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-10T20:12:26.565558+0000) Sep 14 09:45:21 partnach conmon[1735]: 2026-09-14T07:45:21.224+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 18814 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-10T20:12:26.565558+0000) Sep 14 09:45:26 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 19042 slow ops, oldest is log(1000 entries from seq 1002 at 2026-09-10T14:13:26.625340+0000) Sep 14 09:45:26 partnach conmon[1735]: 2026-09-14T07:45:26.700+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 19042 slow ops, oldest is log(1000 entries from seq 1002 at 2026-09-10T14:13:26.625340+0000) Sep 14 09:45:32 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 19280 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-13T08:23:50.558481+0000) Sep 14 09:45:32 partnach conmon[1735]: 2026-09-14T07:45:32.432+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 19280 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-13T08:23:50.558481+0000) Sep 14 09:45:37 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 19433 slow ops, oldest is log(1000 entries from seq 1002 at 2026-09-11T20:28:14.104877+0000) Sep 14 09:45:37 partnach conmon[1735]: 2026-09-14T07:45:37.568+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 19433 slow ops, oldest is log(1000 entries from seq 1002 at 2026-09-11T20:28:14.104877+0000) Sep 14 09:45:43 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 19632 slow ops, oldest is log(1000 entries from seq 2 at 2026-09-10T08:47:13.259138+0000) Sep 14 09:45:43 partnach conmon[1735]: 2026-09-14T07:45:43.292+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 19632 slow ops, oldest is log(1000 entries from seq 2 at 2026-09-10T08:47:13.259138+0000) ``` I’m not quite sure what the lock is here. The mon tries to store some old cluster log entries from various dates? Checked the underlying network fabric, CPU load, clock skew, compacted the rocksdb manually on each mon ect. All is fine. I also set "ceph config set global mon_cluster_log_level info“ **before** the 10th of September. Luckily the upgrade has finished with the MON daemons and is now re-deploying all OSDs. Does anyone know whats going on? Some kind of old hanging Log-Entries? Best Regards, Alex Walender --------------------------- M.Sc Alex Walender Institut für Bio- und Geowissenschaften IBG 5 - Computergestützte Metagenomik / de.NBI Cloud Site Bielefeld Büro : Universität Bielefeld (UHG), N7-101 Tel. : +49-521-106-2907 Forschungszentrum Jülich GmbH 52425 Jülich Sitz der Gesellschaft: Jülich Eingetragen im Handelsregister des Amtsgerichts Düren Nr. HR B 3498 Vorsitzender des Aufsichtsrats: MinDir Stefan Müller Geschäftsführung: Prof. Dr. Astrid Lambrecht (Vorsitzende), Dr. Stephanie Bauer (stellv. Vorsitzende), Prof. Dr. Ir. Pieter Jansens
On 14/09/2026 10:53, Alex Walender wrote:
Hey there!
I’m currently in the process of upgrading the production cluster from v18.2.8 -> v19.2.6 via „ceph orch upgrade“. I tested this upgrade in our staging environment and everything went great. On our production environment I encounter some worrisome issues with the mon daemons.
The cluster itself: 5x MON Daemons 5x MGR Daemons 1500x OSD - 2 larger Pools with each 16k in PG. … so a somewhat larger cluster.
The orchestrator started re-deploying the MGR, which went fine. During the MON redeployment, the cluster went into warning with messages like „34846 slow ops. mon.mon1 has slow ops“. The counter climbed rapidly. Only way to stop this was to restart the MON in question. Sometimes I had to restart the MON multiple times. Afterwards the cluster is healthy again….until I restart any MON. The symptoms return until restarting fixes it again.
I tried to dump the mon_ops_in_flight, but the command just hangs (probably because of the large error count).
Hello, We have a somewhat similar issue on a, busy, CephFS cluster with 1500 OSDs (5 MON+MGR, 7 MDS -- 5 of which are also MON+MGR, 52 OSD nodes, largest pool with 8K PGs). This is an upgrade from 17.2.8 to 19.2.6. We did a staggered upgrade to be able to upgrade the 7 MDS during our next scheduled downtime (with "mgr/orchestrator/fail_fs"), since we have 6 active MDS. The MGR, MON & "crash" upgrades went smoothly, but there were issues during the OSD upgrades (after 24+ hours, about 60% of the OSD updated), and the MONs started to stall: 2266 slow ops, oldest one blocked for 153 sec, mon.$MON0 has slow ops and sometime the MONs did not answer at all: 2026-09-16T12:17:26.485+0000 7f7972084640 -1 monclient: get_monmap_and_config failed to get config Pausing the upgrade seemed to worsen the MONs issues ("159104 slow ops, oldest one blocked for 951 sec, mon.$MON0 has slow ops"), so we resumed it to its completion (41 hours for MGR+MON+crash+1500 OSDs). Sometimes the MON issue goes away on its own, or we have to fail/restart MON(s) to get back to a "normal" situation. This continues even now, after all OSDs have been upgraded & only the MDS (+one RGW) remain in the older version until our scheduled downtime: "overall": { "ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable)": 8, "ceph version 19.2.6 (f9fd95b4335bad6a26d7a74468f55d269e8dbef4) squid (stable)": 1540 } The frequency of the issue is every ~3 hours for this cluster. We see thousands of: "Sep 17 10:25:15 $MON1 ceph-mon[2618906]: ReplicaActive::clear_remote_reservation(): not reserved!" in the MON logs. We had no issue upgrading two smaller clusters (non-staggered upgrades), one with 150 OSDs and another with 800 OSDs (both clusters were already running Squid). Loïc. PS: We find that there are multiple issues with the upgrade process, like: - OSDs in an alternate CRUSH tree were forcibly moved to the "default" branch after their upgrade (on nodes with OSDs in 2 independent CRUSH trees: "default" and "fast" -- no issue on nodes with all OSDs in an alternate CRUSH tree) - the update documentation is lacking regarding the update of the credentials created by "cephadm" itself (bootstrap, admin, deployed services, ...) and, were it not for Eugen B., we probably wouldn't have found "mon_auth_emergency_allowed_ciphers" - progress messages (displayed by "ceph -W cephadm") have a problem with instance counters (all instances of whatever is being updated are displayed as the first instance, there "1/59"), which is a confusing: 2026-09-15T09:11:18.662822+0000 mgr.$MGR1.osnxsd [INF] Upgrade: Updating crash.$NODE7 (1/59) [...] 2026-09-15T09:11:41.166019+0000 mgr.$MGR1.osnxsd [INF] Upgrade: Updating crash.$NODE8 (1/59) (in "src/pybind/mgr/cephadm/upgrade.py", "_upgrade_daemons()": "num" is set to "1" but never incremented in the loop) -- | Loic Tortay <tortay@cc.in2p3.fr> - IN2P3 Computing Centre |
participants (2)
-
Alex Walender
-
Loïc Tortay