Sorry, I forgot to help you with the troubleshooting. Usually, when slow requests appear in the Ceph status output, running 'ceph tell osd.X dump_historic_ops' allows you to see which I/O operations took a long time to execute. The value of X is reported by the Ceph status. Does 'ceph health detail' provide any useful information, such as an OSD ID? Regards, Frédéric. [3] https://docs.ceph.com/en/latest/rados/troubleshooting/troubleshooting-osd/#d... ----- Le 7 Mar 25, à 17:52, Frédéric Nass frederic.nass@univ-lorraine.fr a écrit :
Hi Nicola,
A quick look in Ceph repository points to this commit [1] and this part [2] of the documentation.
I doubt your cluster is actually performing worse than it did before the update. It's just that now it's letting you know when performance isn't optimal (when queries exceed the default threshold).
Regards, Frédéric.
[1] https://github.com/ceph/ceph/commit/73b80a9a2c38259346fb646f85fa2ba4dcbb1329 [2] https://docs.ceph.com/en/latest/rados/operations/health-checks/#block-device...
----- Le 7 Mar 25, à 17:05, Nicola Mori mori@fi.infn.it a écrit :
Dear Ceph users,
after upgrading from 19.2.0 to 19.2.1 (via cephadm) my cluster started showing some warnings never seen before:
29 OSD(s) experiencing slow operations in BlueStore 13 OSD(s) experiencing stalled read in db device of BlueFS
I searched for these messages but didn't find much. I also noticed that when browsing the CephFS folders (using the kernel module for the client) sometimes the client gets stuck for a long time before showing the folder content; however I don't know if this can be related to the above warnings.
I'd need help to understand what's going on and eventually how to troubleshoot it. Thanks,
Nicola
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io