Resolving a pg inconsistent Issue
Hi Ceph users, We are running a suse SES 5.5 cluster that's largely based on luminous with some mimic backports. We've been doing some large reshuffling from adding in additional OSDs and during this process we have an inconsistent pg group, investigation suggests there was a read error. We would like to target it with a deep-scrub and possibly repair attempt before marking the individual osd down and replacing the drive, but because of the weeks of reshuffling, our cluster is behind on deep-scrubs and it seemingly is refusing to scrub the marked pg, or at least it is scheduling other pg for deep-scrubs first. Is there any method we can use to grant deep-scrub priority? How about adjusting deep-scrub timeframes to be 6 months temporarily, would that allow us to force the desired pg deep-scrub to occur prior? Or could set some of the no scrub flags and proceed with a pg repair attempt? Thank you for any suggestions or advice, -- Steven Pine webair.com *P* 516.938.4100 x *E * steven.pine@webair.com <https://www.facebook.com/WebairInc/> <https://www.linkedin.com/company/webair>
Hi, I'm not sure what SUSE Support would suggest (you probably should be able to open a case?) but I'd probably go for nodeep-scub flag, wait for the already started deep-scrubs to finish and then trigger a pg repair and a deep-scrub on the pg if the repair doesn't resolve it. This shouldn't take too long, afterwards unset the flag. I don't know if there are other/better ways to prioritize a deep-scrub. Regards, Eugen Zitat von Steven Pine <steven.pine@webair.com>:
Hi Ceph users,
We are running a suse SES 5.5 cluster that's largely based on luminous with some mimic backports.
We've been doing some large reshuffling from adding in additional OSDs and during this process we have an inconsistent pg group, investigation suggests there was a read error.
We would like to target it with a deep-scrub and possibly repair attempt before marking the individual osd down and replacing the drive, but because of the weeks of reshuffling, our cluster is behind on deep-scrubs and it seemingly is refusing to scrub the marked pg, or at least it is scheduling other pg for deep-scrubs first.
Is there any method we can use to grant deep-scrub priority? How about adjusting deep-scrub timeframes to be 6 months temporarily, would that allow us to force the desired pg deep-scrub to occur prior? Or could set some of the no scrub flags and proceed with a pg repair attempt?
Thank you for any suggestions or advice,
-- Steven Pine webair.com *P* 516.938.4100 x *E * steven.pine@webair.com
<https://www.facebook.com/WebairInc/> <https://www.linkedin.com/company/webair> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (2)
-
Eugen Block
-
Steven Pine