12 Nov
2023
12 Nov
'23
11:28 p.m.
during scrubbing, OSD latency spikes to 300-600 ms,
I have seen Ceph clusters spike to several seconds per IO operation as they were designed for the same goals.
resulting in sluggish performance for all VMs. Additionally, some OSDs fail during the scrubbing process.
Most likely they time out because of IO congestion rather than failing.
bluestore(/var/lib/ceph/osd/ceph-10) log_latency slow operation observed for next, latency = 74835459564ns bluestore(/var/lib/ceph/osd/ceph-10) log_latency slow operation observed for next, latency = 42822161884ns
7.48s? 4.28s?