On Sun, Apr 25, 2021 at 12:37 AM Markus Kienast <mark@trickkiste.at> wrote:
I am seeing these messages when booting from RBD and booting hangs there.
libceph: get_reply osd2 tid 1459933 data 3248128 > preallocated 131072, skipping
However, Ceph Health is OK, so I have no idea what is going on. I reboot my 3 node cluster and it works again for about two weeks.
How can I find out more about this issue, how can I dig deeper? Also there has been at least one report about this issue before on this mailing list - "[ceph-users] Strange Data Issue - Unexpected client hang on OSD I/O Error" - but no solution has been presented.
This report was from 2018, so no idea if this is still an issue for Dyweni the original reporter. If you read this, I would be happy to hear how you solved the problem.
Hi Markus, What versions of ceph and the kernel are in use? Are you also seeing I/O errors and "missing primary copy of ..., will try copies on ..." messages in the OSD logs (in this case osd2)? Thanks, Ilya