I have a Nautilus cluster with 7 nodes, 210 HDDs. I recently added the 7th node with 30 OSDs which are currently rebalancing very slowly. I just noticed that the ethernet interface only negotiated a 1Gb connection, even though it has a 10Gb interface. I’m not sure why, but would like to reboot the node to get the interface back to 10Gb. Is it ok to do this? What should I do to prep the cluster for the reboot? Jeffrey Turmelle International Research Institute for Climate & Society <https://iri.columbia.edu/> The Climate School <https://climate.columbia.edu/> at Columbia University <https://columbia.edu/> 845-652-3461
Hi, if your failure domain is "host" and you have enough redundancy (e.g. replicated size 3 or proper erasure-code profiles and rulesets) you should be able to reboot without any issue. Depending on how long the reboot would take, you could set the noout flag, the default are 10 minutes until OSDs are marked "out". Regards, Eugen Zitat von Jeffrey Turmelle <jefft@iri.columbia.edu>:
I have a Nautilus cluster with 7 nodes, 210 HDDs. I recently added the 7th node with 30 OSDs which are currently rebalancing very slowly. I just noticed that the ethernet interface only negotiated a 1Gb connection, even though it has a 10Gb interface. I’m not sure why, but would like to reboot the node to get the interface back to 10Gb.
Is it ok to do this? What should I do to prep the cluster for the reboot?
Jeffrey Turmelle International Research Institute for Climate & Society <https://iri.columbia.edu/> The Climate School <https://climate.columbia.edu/> at Columbia University <https://columbia.edu/> 845-652-3461
Den tors 2 mars 2023 kl 08:09 skrev Eugen Block <eblock@nde.ag>:
if your failure domain is "host" and you have enough redundancy (e.g. replicated size 3 or proper erasure-code profiles and rulesets) you should be able to reboot without any issue. Depending on how long the reboot would take, you could set the noout flag, the default are 10 minutes until OSDs are marked "out".
"ceph osd set noout" if you need more than 10 minutes. The other copies of the now-missing PGs will handle it in the meantime while you restart. -- May the most significant bit of your life be positive.
Hey Jeff, As long as you set the maintenance flags (noout/norebalance) you should be good to take the node down with a reboot Regards, Bailey
From: Jeffrey Turmelle <jefft@iri.columbia.edu> Sent: March 1, 2023 2:47 PM To: ceph-users@ceph.io Subject: [ceph-users] Interruption of rebalancing
I have a Nautilus cluster with 7 nodes, 210 HDDs. I recently added the 7th node with 30 OSDs which are currently rebalancing very slowly. I just noticed that >the ethernet interface only negotiated a 1Gb connection, even though it has a 10Gb interface. I’m not sure why, but would like to reboot the node to get the >interface back to 10Gb.
Is it ok to do this? What should I do to prep the cluster for the reboot?
Jeffrey Turmelle
International Research Institute for Climate <https://iri.columbia.edu> & Society
The Climate School <https://climate.columbia.edu> at Columbia University <https://columbia.edu>
845-652-3461
Thanks everyone for the help. I set noout on the cluster, rebooted the node and it came back to rebalancing/remapping where it left off. CEPH is fantastic.
From: Jeffrey Turmelle <jefft@iri.columbia.edu <mailto:jefft@iri.columbia.edu>> Sent: March 1, 2023 2:47 PM To: ceph-users@ceph.io <mailto:ceph-users@ceph.io> Subject: [ceph-users] Interruption of rebalancing
I have a Nautilus cluster with 7 nodes, 210 HDDs. I recently added the 7th node with 30 OSDs which are currently rebalancing very slowly. I just noticed that >the ethernet interface only negotiated a 1Gb connection, even though it has a 10Gb interface. I’m not sure why, but would like to reboot the node to get the >interface back to 10Gb.
Is it ok to do this? What should I do to prep the cluster for the reboot?
Jeffrey Turmelle International Research Institute for Climate & Society <https://iri.columbia.edu/> The Climate School <https://climate.columbia.edu/> at Columbia University <https://columbia.edu/> 845-652-3461
participants (4)
-
Bailey Allison
-
Eugen Block
-
Janne Johansson
-
Jeffrey Turmelle