Slow recovery and inaccurate recovery figures since Quincy upgrade
Hi, We run 3 production clusters in a multi-site setup. They were deployed with Ceph-Ansible but recently switched to cephadm while on the Pacific release. Shortly after migrating to cephadm they were upgraded to Quincy. Since moving to Quincy, the recovery on one of the replica sites has tanked quite severely. We use 8Tb SAS HDDs for OSDs with WAL and DB on nVME drives. Before the upgrade it would take 1-2 days to resilver a OSD, but recently we replaced a drive and it took 6 days to resilver. Nothing else has changed with the configuration of the cluster, and the other 2 are performing recovery fine as far as we can see. Does anyone have any ideas of what could be the issue here or anywhere we can check what is going on?? Also while recovering, we've noticed inaccurate recovery information in ceph -s. We've seen the recovery section of ceph -s reporting its running at less than 10 keys/s, but looking at the Degraded data redundancy we see this drop multiple 100s of keys per second. Does anyone have any advice they can offer on this too please?? Cheers Iain [ceph: root@gb4-li-cephgw-001 /]# ceph -s; sleep 3; ceph -s cluster: id: 6dabcf41-90d7-4e90-b259-1cc0bf298052 health: HEALTH_WARN noout,norebalance flag(s) set Degraded data redundancy: 59610/371260626 objects degraded (0.016%), 1 pg degraded, 1 pg undersized services: mon: 3 daemons, quorum gb4-li-cephgw-001,gb4-li-cephgw-002,gb4-li-cephgw-003 (age 2h) mgr: gb4-li-cephgw-003(active, since 2h), standbys: gb4-li-cephgw-002.iqmxgu, gb4-li-cephgw-001 osd: 72 osds: 72 up (since 39m), 72 in (since 3d); 1 remapped pgs flags noout,norebalance rgw: 3 daemons active (3 hosts, 1 zones) data: pools: 11 pools, 1457 pgs objects: 63.76M objects, 173 TiB usage: 251 TiB used, 275 TiB / 526 TiB avail pgs: 59610/371260626 objects degraded (0.016%) 1452 active+clean 4 active+clean+scrubbing+deep 1 active+undersized+degraded+remapped+backfilling io: client: 249 MiB/s rd, 611 KiB/s wr, 355 op/s rd, 402 op/s wr recovery: 13 KiB/s, 3 keys/s, 2 objects/s progress: Global Recovery Event (39m) [===========================.] (remaining: 1s) cluster: id: 6dabcf41-90d7-4e90-b259-1cc0bf298052 health: HEALTH_WARN noout,norebalance flag(s) set Degraded data redundancy: 59116/371260644 objects degraded (0.016%), 1 pg degraded, 1 pg undersized services: mon: 3 daemons, quorum gb4-li-cephgw-001,gb4-li-cephgw-002,gb4-li-cephgw-003 (age 2h) mgr: gb4-li-cephgw-003(active, since 2h), standbys: gb4-li-cephgw-002.iqmxgu, gb4-li-cephgw-001 osd: 72 osds: 72 up (since 39m), 72 in (since 3d); 1 remapped pgs flags noout,norebalance rgw: 3 daemons active (3 hosts, 1 zones) data: pools: 11 pools, 1457 pgs objects: 63.76M objects, 173 TiB usage: 251 TiB used, 275 TiB / 526 TiB avail pgs: 59116/371260644 objects degraded (0.016%) 1452 active+clean 4 active+clean+scrubbing+deep 1 active+undersized+degraded+remapped+backfilling io: client: 258 MiB/s rd, 595 KiB/s wr, 346 op/s rd, 387 op/s wr recovery: 15 KiB/s, 2 keys/s, 2 objects/s progress: Global Recovery Event (39m) [===========================.] (remaining: 1s) [ceph: root@gb4-li-cephgw-001 /]# ceph -s; sleep 3; ceph -s cluster: id: 6dabcf41-90d7-4e90-b259-1cc0bf298052 health: HEALTH_WARN noout,norebalance flag(s) set Degraded data redundancy: 58503/371260638 objects degraded (0.016%), 1 pg degraded, 1 pg undersized services: mon: 3 daemons, quorum gb4-li-cephgw-001,gb4-li-cephgw-002,gb4-li-cephgw-003 (age 2h) mgr: gb4-li-cephgw-003(active, since 2h), standbys: gb4-li-cephgw-002.iqmxgu, gb4-li-cephgw-001 osd: 72 osds: 72 up (since 39m), 72 in (since 3d); 1 remapped pgs flags noout,norebalance rgw: 3 daemons active (3 hosts, 1 zones) data: pools: 11 pools, 1457 pgs objects: 63.76M objects, 173 TiB usage: 251 TiB used, 275 TiB / 526 TiB avail pgs: 58503/371260638 objects degraded (0.016%) 1452 active+clean 4 active+clean+scrubbing+deep 1 active+undersized+degraded+remapped+backfilling io: client: 245 MiB/s rd, 278 KiB/s wr, 247 op/s rd, 183 op/s wr recovery: 16 KiB/s, 2 keys/s, 2 objects/s progress: Global Recovery Event (39m) [===========================.] (remaining: 1s) cluster: id: 6dabcf41-90d7-4e90-b259-1cc0bf298052 health: HEALTH_WARN noout,norebalance flag(s) set Degraded data redundancy: 58157/371260644 objects degraded (0.016%), 1 pg degraded, 1 pg undersized services: mon: 3 daemons, quorum gb4-li-cephgw-001,gb4-li-cephgw-002,gb4-li-cephgw-003 (age 2h) mgr: gb4-li-cephgw-003(active, since 2h), standbys: gb4-li-cephgw-002.iqmxgu, gb4-li-cephgw-001 osd: 72 osds: 72 up (since 39m), 72 in (since 3d); 1 remapped pgs flags noout,norebalance rgw: 3 daemons active (3 hosts, 1 zones) data: pools: 11 pools, 1457 pgs objects: 63.76M objects, 173 TiB usage: 251 TiB used, 275 TiB / 526 TiB avail pgs: 58157/371260644 objects degraded (0.016%) 1452 active+clean 4 active+clean+scrubbing+deep 1 active+undersized+degraded+remapped+backfilling io: client: 243 MiB/s rd, 285 KiB/s wr, 252 op/s rd, 197 op/s wr recovery: 13 KiB/s, 0 keys/s, 1 objects/s progress: Global Recovery Event (39m) [===========================.] (remaining: 1s) [ceph: root@gb4-li-cephgw-001 /]# Iain Stott OpenStack Engineer Iain.Stott@thehutgroup.com [THG Ingenuity Logo]<https://www.thg.com> [https://i.imgur.com/wbpVRW6.png]<https://www.linkedin.com/company/thgplc/?originalSubdomain=uk> [https://i.imgur.com/c3040tr.png] <https://twitter.com/thgplc?lang=en>
Hello Iain, Does anyone have any ideas of what could be the issue here or anywhere we
can check what is going on??
You could be hitting the slow backfill/recovery issue with mclock_scheduler. Could you please provide the output of the following commands? 1. ceph versions 2. ceph config get osd.<id> osd_op_queue 3. ceph config show osd.<id> | grep osd_max_backfills 4. ceph config show osd.<id> | grep osd_recovery_max_active 5. ceph config show-with-defaults osd.<id> | grep osd_mclock where 'id' can be any valid osd id With the mclock_scheduler enabled and with 17.2.5, it is not possible to override recovery settings like 'osd_max_backfills' and other recovery related config options. To improve the recovery rate, you can temporarily switch the mClock profile to 'high_recovery_ops' on all the OSDs by issuing: ceph config set osd osd_mclock_profile high_recovery_ops During recovery with this profile, you may notice a dip in the client ops performance which is expected. Once the recovery is done, you can switch the mClock profile back to the default 'high_client_ops' profile. Please note that the upcoming Quincy release will address the slow backfill issues along with other usability improvements. -Sridhar
Hi Sridhar, Thanks for the response, I have added the output you requested below, I have attached the output from the last command in a file as it was rather long. We did try to set high_recovery_ops but it didn't seem to have any visible effect. root@gb4-li-cephgw-001 ~ # ceph versions { "mon": { "ceph version 17.2.6 (d7ff0d10654d2280e08f1ab989c7cdf3064446a5) quincy (stable)": 3 }, "mgr": { "ceph version 17.2.6 (d7ff0d10654d2280e08f1ab989c7cdf3064446a5) quincy (stable)": 3 }, "osd": { "ceph version 17.2.6 (d7ff0d10654d2280e08f1ab989c7cdf3064446a5) quincy (stable)": 72 }, "mds": {}, "rgw": { "ceph version 17.2.6 (d7ff0d10654d2280e08f1ab989c7cdf3064446a5) quincy (stable)": 3 }, "overall": { "ceph version 17.2.6 (d7ff0d10654d2280e08f1ab989c7cdf3064446a5) quincy (stable)": 81 } } root@gb4-li-cephgw-001 ~ # ceph config get osd.0 osd_op_queue mclock_scheduler root@gb4-li-cephgw-001 ~ # ceph config show osd.0 | grep osd_max_backfills osd_max_backfills 3 override (mon[3]),default[10] root@gb4-li-cephgw-001 ~ # ceph config show osd.0 | grep osd_recovery_max_active osd_recovery_max_active 9 override (mon[9]),default[0] osd_recovery_max_active_hdd 10 default osd_recovery_max_active_ssd 20 default Thanks Iain ________________________________ From: Sridhar Seshasayee <sseshasa@redhat.com> Sent: 03 October 2023 09:07 To: Iain Stott <Iain.Stott@thehutgroup.com> Cc: ceph-users@ceph.io <ceph-users@ceph.io>; dl-osadmins <dl-osadmins@thehutgroup.com> Subject: Re: [ceph-users] Slow recovery and inaccurate recovery figures since Quincy upgrade CAUTION: This email originates from outside THG ________________________________ Hello Iain, Does anyone have any ideas of what could be the issue here or anywhere we can check what is going on?? You could be hitting the slow backfill/recovery issue with mclock_scheduler. Could you please provide the output of the following commands? 1. ceph versions 2. ceph config get osd.<id> osd_op_queue 3. ceph config show osd.<id> | grep osd_max_backfills 4. ceph config show osd.<id> | grep osd_recovery_max_active 5. ceph config show-with-defaults osd.<id> | grep osd_mclock where 'id' can be any valid osd id With the mclock_scheduler enabled and with 17.2.5, it is not possible to override recovery settings like 'osd_max_backfills' and other recovery related config options. To improve the recovery rate, you can temporarily switch the mClock profile to 'high_recovery_ops' on all the OSDs by issuing: ceph config set osd osd_mclock_profile high_recovery_ops During recovery with this profile, you may notice a dip in the client ops performance which is expected. Once the recovery is done, you can switch the mClock profile back to the default 'high_client_ops' profile. Please note that the upcoming Quincy release will address the slow backfill issues along with other usability improvements. -Sridhar
To help complete the recovery, you can temporarily try disabling scrub and deep scrub operations by running: ceph osd set noscrub ceph osd set nodeep-scrub This should help speed up the recovery process. Once the recovery is done, you can unset the above scrub flags and revert the mClock profile back to 'high_client_ops'. If the above doesn't help, then there's something else that's causing the slow recovery. -Sridhar
participants (3)
-
Iain Stott
-
Sake
-
Sridhar Seshasayee