Hey there! I’m currently in the process of upgrading the production cluster from v18.2.8 -> v19.2.6 via „ceph orch upgrade“. I tested this upgrade in our staging environment and everything went great. On our production environment I encounter some worrisome issues with the mon daemons. The cluster itself: 5x MON Daemons 5x MGR Daemons 1500x OSD - 2 larger Pools with each 16k in PG. … so a somewhat larger cluster. The orchestrator started re-deploying the MGR, which went fine. During the MON redeployment, the cluster went into warning with messages like „34846 slow ops. mon.mon1 has slow ops“. The counter climbed rapidly. Only way to stop this was to restart the MON in question. Sometimes I had to restart the MON multiple times. Afterwards the cluster is healthy again….until I restart any MON. The symptoms return until restarting fixes it again. I tried to dump the mon_ops_in_flight, but the command just hangs (probably because of the large error count). Filtering the journal for the issues: ``` journalctl -g "oldest is" --since "3 hours ago" -u ceph-d8a6f365-99de-42f2-b32b-c9ea11df945c@mon.partnach <mailto:ceph-d8a6f365-99de-42f2-b32b-c9ea11df945c@mon.partnach> Sep 14 09:45:05 partnach conmon[1735]: 2026-09-14T07:45:05.331+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 18286 slow ops, oldest is log(1000 entries from seq 2 at 2026-08-26T04:51:47.819315+0000) Sep 14 09:45:10 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 18505 slow ops, oldest is log(1000 entries from seq 4003 at 2026-09-11T12:25:52.136679+0000) Sep 14 09:45:10 partnach conmon[1735]: 2026-09-14T07:45:10.827+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 18505 slow ops, oldest is log(1000 entries from seq 4003 at 2026-09-11T12:25:52.136679+0000) Sep 14 09:45:15 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 18671 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-07T16:47:58.529125+0000) Sep 14 09:45:15 partnach conmon[1735]: 2026-09-14T07:45:15.924+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 18671 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-07T16:47:58.529125+0000) Sep 14 09:45:21 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 18814 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-10T20:12:26.565558+0000) Sep 14 09:45:21 partnach conmon[1735]: 2026-09-14T07:45:21.224+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 18814 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-10T20:12:26.565558+0000) Sep 14 09:45:26 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 19042 slow ops, oldest is log(1000 entries from seq 1002 at 2026-09-10T14:13:26.625340+0000) Sep 14 09:45:26 partnach conmon[1735]: 2026-09-14T07:45:26.700+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 19042 slow ops, oldest is log(1000 entries from seq 1002 at 2026-09-10T14:13:26.625340+0000) Sep 14 09:45:32 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 19280 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-13T08:23:50.558481+0000) Sep 14 09:45:32 partnach conmon[1735]: 2026-09-14T07:45:32.432+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 19280 slow ops, oldest is log(1000 entries from seq 2002 at 2026-09-13T08:23:50.558481+0000) Sep 14 09:45:37 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 19433 slow ops, oldest is log(1000 entries from seq 1002 at 2026-09-11T20:28:14.104877+0000) Sep 14 09:45:37 partnach conmon[1735]: 2026-09-14T07:45:37.568+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 19433 slow ops, oldest is log(1000 entries from seq 1002 at 2026-09-11T20:28:14.104877+0000) Sep 14 09:45:43 partnach ceph-mon[1742]: mon.partnach@0(leader) e78 get_health_metrics reporting 19632 slow ops, oldest is log(1000 entries from seq 2 at 2026-09-10T08:47:13.259138+0000) Sep 14 09:45:43 partnach conmon[1735]: 2026-09-14T07:45:43.292+0000 7fe90bfaf640 -1 mon.partnach@0(leader) e78 get_health_metrics reporting 19632 slow ops, oldest is log(1000 entries from seq 2 at 2026-09-10T08:47:13.259138+0000) ``` I’m not quite sure what the lock is here. The mon tries to store some old cluster log entries from various dates? Checked the underlying network fabric, CPU load, clock skew, compacted the rocksdb manually on each mon ect. All is fine. I also set "ceph config set global mon_cluster_log_level info“ **before** the 10th of September. Luckily the upgrade has finished with the MON daemons and is now re-deploying all OSDs. Does anyone know whats going on? Some kind of old hanging Log-Entries? Best Regards, Alex Walender --------------------------- M.Sc Alex Walender Institut für Bio- und Geowissenschaften IBG 5 - Computergestützte Metagenomik / de.NBI Cloud Site Bielefeld Büro : Universität Bielefeld (UHG), N7-101 Tel. : +49-521-106-2907 Forschungszentrum Jülich GmbH 52425 Jülich Sitz der Gesellschaft: Jülich Eingetragen im Handelsregister des Amtsgerichts Düren Nr. HR B 3498 Vorsitzender des Aufsichtsrats: MinDir Stefan Müller Geschäftsführung: Prof. Dr. Astrid Lambrecht (Vorsitzende), Dr. Stephanie Bauer (stellv. Vorsitzende), Prof. Dr. Ir. Pieter Jansens