Exit yolo mode by increasing size/min_size does not (really) work
Hi! 😊 It would be very kind of you to help us with that! We have pools in our ceph cluster that are set to replicated size 2 min_size 1. Obviously we want to go to size 3 / min_size 2 but we experience problems with that. USED goes to 100% instantly and MAX AVAIL goes to 0. Write operations seemed to stop. POOLS: NAME ID USED %USED MAX AVAIL OBJECTS Pool1 24 35791G 35.04 66339G 8927762 Pool2 25 11610G 14.89 66339G 3004740 Pool3 26 17557G 100.00 0 2666972 Before the change it was like this: NAME ID USED %USED MAX AVAIL OBJECTS Pool1 24 35791G 35.04 66339G 8927762 Pool2 25 11610G 14.89 66339G 3004740 Pool3 26 17558G 20.93 66339G 2667013 This was quite surprising to us as we’d expect USED to go to something like 30%. Going back to 2/1 also gave us back the 20.93% usage instantly. What’s the matter here? Thank you and best regards Stefan ________________________________ BearingPoint GmbH Sitz: Wien Firmenbuchgericht: Handelsgericht Wien Firmenbuchnummer: FN 175524z The information in this email is confidential and may be legally privileged. If you are not the intended recipient of this message, any review, disclosure, copying, distribution, retention, or any action taken or omitted to be taken in reliance on it is prohibited and may be unlawful. If you are not the intended recipient, please reply to or forward a copy of this message to the sender and delete the message, any attachments, and any copies thereof from your system.
Hi, I don't have an explanation yet but some more information about your cluster would be useful like 'ceph osd tree', 'ceph osd df', 'ceph status' etc. Thanks, Eugen Zitat von Stefan Pinter <stefan.pinter@bearingpoint.com>:
Hi! 😊
It would be very kind of you to help us with that!
We have pools in our ceph cluster that are set to replicated size 2 min_size 1. Obviously we want to go to size 3 / min_size 2 but we experience problems with that.
USED goes to 100% instantly and MAX AVAIL goes to 0. Write operations seemed to stop.
POOLS: NAME ID USED %USED MAX AVAIL OBJECTS Pool1 24 35791G 35.04 66339G 8927762 Pool2 25 11610G 14.89 66339G 3004740 Pool3 26 17557G 100.00 0 2666972
Before the change it was like this:
NAME ID USED %USED MAX AVAIL OBJECTS Pool1 24 35791G 35.04 66339G 8927762 Pool2 25 11610G 14.89 66339G 3004740 Pool3 26 17558G 20.93 66339G 2667013
This was quite surprising to us as we’d expect USED to go to something like 30%. Going back to 2/1 also gave us back the 20.93% usage instantly.
What’s the matter here?
Thank you and best regards Stefan ________________________________ BearingPoint GmbH Sitz: Wien Firmenbuchgericht: Handelsgericht Wien Firmenbuchnummer: FN 175524z
The information in this email is confidential and may be legally privileged. If you are not the intended recipient of this message, any review, disclosure, copying, distribution, retention, or any action taken or omitted to be taken in reliance on it is prohibited and may be unlawful. If you are not the intended recipient, please reply to or forward a copy of this message to the sender and delete the message, any attachments, and any copies thereof from your system. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
hi, thank you Eugen for being interested in solving this ;) certainly, here are some more infos: ceph osd tree https://privatebin.net/?db7b93a623095879#AKJNy6pKNxa5XssjUpxxjMnggc3d4PirTH1... ceph osd df https://privatebin.net/?0f7c3b091b683d65#8K4KQW5a2G2mFgcnTdUjQXJvcZCpAJGxcPR... ceph status https://privatebin.net/?e49d305a69d022cd#FtAXq8ZoMqmGa67UFdv6k1xaCqRqmgL5FXc...
Could you also share more details about the pools: ceph osd pool ls detail ceph osd crush rule dump <RULE> (for each pool if they use different rules) Thanks Eugen Zitat von stefan.pinter@bearingpoint.com:
hi, thank you Eugen for being interested in solving this ;)
certainly, here are some more infos:
ceph osd tree https://privatebin.net/?db7b93a623095879#AKJNy6pKNxa5XssjUpxxjMnggc3d4PirTH1...
ceph osd df https://privatebin.net/?0f7c3b091b683d65#8K4KQW5a2G2mFgcnTdUjQXJvcZCpAJGxcPR...
ceph status https://privatebin.net/?e49d305a69d022cd#FtAXq8ZoMqmGa67UFdv6k1xaCqRqmgL5FXc... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
sure! ceph osd pool ls detail https://privatebin.net/?85105578dd50f65f#4oNunvNfLoNbnqJwuXoWXrB1idt4zMGnBXd... i guess this needs some cleaning up regarding snapshots - could this be a problem? ceph osd crush rule dump https://privatebin.net/?bd589bc9d7800dd3#3PFS3659qXqbxfaXSUcKot3ynmwRG2mDjpx...
Okay, so your applied crush rule has failure domain „room“ which you have three of, but the third has no OSDs available. Check your osd tree output, that’s why ceph fails to create a third replica. To resolve this you can either change the rule to a different failure domain (for example „host“) and then increase the size. Or you create a new rule and apply it to the pool(s). Either way you’ll have to decide how to place three replicas, e. g. move hosts within the crush tree (and probably the third room into the default root) to enable an even distribution. Note that moving buckets within the crush tree will cause rebalancing. Regards, Eugen Zitat von stefan.pinter@bearingpoint.com:
sure!
ceph osd pool ls detail https://privatebin.net/?85105578dd50f65f#4oNunvNfLoNbnqJwuXoWXrB1idt4zMGnBXd... i guess this needs some cleaning up regarding snapshots - could this be a problem?
ceph osd crush rule dump https://privatebin.net/?bd589bc9d7800dd3#3PFS3659qXqbxfaXSUcKot3ynmwRG2mDjpx... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (3)
-
Eugen Block
-
Stefan Pinter
-
stefan.pinterï¼ bearingpoint.com