Upgrading nautilus / centos7 to octopus / ubuntu 20.04. - Suggestions and hints?
Hi, As I’v read and thought a lot about the migration as this is a bigger project, I was wondering if anyone has done that already and might share some notes or playbooks, because in all readings there where some parts missing or miss understandable to me. I do have some different approaches in mind, so may be you have some suggestions or hints. a) upgrade nautilus on centos 7 with the few missing features like dashboard and prometheus. After that migrate one node after an other to ubuntu 20.04 with octopus and than upgrade ceph to the recent stable version. b) migrate one node after an other to ubuntu 18.04 with nautilus and then upgrade to octupus and after that to ubuntu 20.04. or c) upgrade one node after an other to ubuntu 20.04 with octopus and join it to the cluster until all nodes are upgraded. For test I tried c) with a mon node, but adding that to the cluster fails with some failed state, still probing for the other mons. (I dont have the right log at hand right now.) So my questions are: a) What would be the best (most stable) migration path and b) is it in general possible to add a new octopus mon (not upgraded one) to a nautilus cluster, where the other mons are still on nautilus? I hope my thoughts and questions are understandable :) Thanks for any hint and suggestion. Best . Götz
Hi Goetz, I've done the same, and went to Octopus and to Ubuntu. It worked like a charm and with pip, you can get the pecan library working. I think I did it with this: yum -y install python36-six.noarch python36-PyYAML.x86_64 pip3 install pecan werkzeug cherrypy Worked very well, until we got hit by this bug: https://tracker.ceph.com/issues/53729#note-65 Nautilus seem not to have tooling to detect it, and the fix is not backported to octopus. And because our clusters started to act badly after the octopus upgrade, and we fast forwarded to pacific (untested emergency cluster upgrades are okayish but ugly :D ). And because of the bug, we went another route with the last cluster. I reinstalled all hosts with ubuntu 18.04, then update straight to pacific, and then upgrade to ubuntu 20.04. Hope that helped. Cheers Boris Am Di., 1. Aug. 2023 um 20:06 Uhr schrieb Götz Reinicke < goetz.reinicke@filmakademie.de>:
Hi,
As I’v read and thought a lot about the migration as this is a bigger project, I was wondering if anyone has done that already and might share some notes or playbooks, because in all readings there where some parts missing or miss understandable to me.
I do have some different approaches in mind, so may be you have some suggestions or hints.
a) upgrade nautilus on centos 7 with the few missing features like dashboard and prometheus. After that migrate one node after an other to ubuntu 20.04 with octopus and than upgrade ceph to the recent stable version.
b) migrate one node after an other to ubuntu 18.04 with nautilus and then upgrade to octupus and after that to ubuntu 20.04.
or
c) upgrade one node after an other to ubuntu 20.04 with octopus and join it to the cluster until all nodes are upgraded.
For test I tried c) with a mon node, but adding that to the cluster fails with some failed state, still probing for the other mons. (I dont have the right log at hand right now.)
So my questions are:
a) What would be the best (most stable) migration path and
b) is it in general possible to add a new octopus mon (not upgraded one) to a nautilus cluster, where the other mons are still on nautilus?
I hope my thoughts and questions are understandable :)
Thanks for any hint and suggestion. Best . Götz _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal.
Hi Goetz, Which method you finally choose? We've done a successful migration from Centos 8 to ubuntu 20.04 but we have a centos 7 nautilus cluster which we'd like to move to Ubuntu 20.04 octopus same as you. Wonder any of you tried to skip Rocky 8 from the flow? Thank you ________________________________ From: Boris Behrens <bb@kervyn.de> Sent: Wednesday, August 2, 2023 1:24 AM To: Götz Reinicke <goetz.reinicke@filmakademie.de> Cc: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [ceph-users] Re: Upgrading nautilus / centos7 to octopus / ubuntu 20.04. - Suggestions and hints? Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________ Hi Goetz, I've done the same, and went to Octopus and to Ubuntu. It worked like a charm and with pip, you can get the pecan library working. I think I did it with this: yum -y install python36-six.noarch python36-PyYAML.x86_64 pip3 install pecan werkzeug cherrypy Worked very well, until we got hit by this bug: https://tracker.ceph.com/issues/53729#note-65 Nautilus seem not to have tooling to detect it, and the fix is not backported to octopus. And because our clusters started to act badly after the octopus upgrade, and we fast forwarded to pacific (untested emergency cluster upgrades are okayish but ugly :D ). And because of the bug, we went another route with the last cluster. I reinstalled all hosts with ubuntu 18.04, then update straight to pacific, and then upgrade to ubuntu 20.04. Hope that helped. Cheers Boris Am Di., 1. Aug. 2023 um 20:06 Uhr schrieb Götz Reinicke < goetz.reinicke@filmakademie.de>:
Hi,
As I’v read and thought a lot about the migration as this is a bigger project, I was wondering if anyone has done that already and might share some notes or playbooks, because in all readings there where some parts missing or miss understandable to me.
I do have some different approaches in mind, so may be you have some suggestions or hints.
a) upgrade nautilus on centos 7 with the few missing features like dashboard and prometheus. After that migrate one node after an other to ubuntu 20.04 with octopus and than upgrade ceph to the recent stable version.
b) migrate one node after an other to ubuntu 18.04 with nautilus and then upgrade to octupus and after that to ubuntu 20.04.
or
c) upgrade one node after an other to ubuntu 20.04 with octopus and join it to the cluster until all nodes are upgraded.
For test I tried c) with a mon node, but adding that to the cluster fails with some failed state, still probing for the other mons. (I dont have the right log at hand right now.)
So my questions are:
a) What would be the best (most stable) migration path and
b) is it in general possible to add a new octopus mon (not upgraded one) to a nautilus cluster, where the other mons are still on nautilus?
I hope my thoughts and questions are understandable :)
Thanks for any hint and suggestion. Best . Götz _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io ________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses.
I have compiled nautilus for el9 and am going to test adding a el9 osd node the the existing el7 cluster. If that is ok, I will upgrade all nodes first to el9.
-----Original Message----- From: Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com> Sent: Wednesday, 17 January 2024 08:09 To: ballison@45drives.com; Eugen Block <eblock@nde.ag>; Götz Reinicke <goetz.reinicke@filmakademie.de>; Boris Behrens <bb@kervyn.de> Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Upgrading nautilus / centos7 to octopus / ubuntu 20.04. - Suggestions and hints?
Hi Goetz,
Which method you finally choose? We've done a successful migration from Centos 8 to ubuntu 20.04 but we have a centos 7 nautilus cluster which we'd like to move to Ubuntu 20.04 octopus same as you.
Wonder any of you tried to skip Rocky 8 from the flow?
Thank you
________________________________ From: Boris Behrens <bb@kervyn.de> Sent: Wednesday, August 2, 2023 1:24 AM To: Götz Reinicke <goetz.reinicke@filmakademie.de> Cc: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [ceph-users] Re: Upgrading nautilus / centos7 to octopus / ubuntu 20.04. - Suggestions and hints?
Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________
Hi Goetz, I've done the same, and went to Octopus and to Ubuntu. It worked like a charm and with pip, you can get the pecan library working. I think I did it with this: yum -y install python36-six.noarch python36-PyYAML.x86_64 pip3 install pecan werkzeug cherrypy
Worked very well, until we got hit by this bug: https://tracker.ceph.com/issues/53729#note-65 Nautilus seem not to have tooling to detect it, and the fix is not backported to octopus.
And because our clusters started to act badly after the octopus upgrade, and we fast forwarded to pacific (untested emergency cluster upgrades are okayish but ugly :D ).
And because of the bug, we went another route with the last cluster. I reinstalled all hosts with ubuntu 18.04, then update straight to pacific, and then upgrade to ubuntu 20.04.
Hope that helped.
Cheers Boris
Am Di., 1. Aug. 2023 um 20:06 Uhr schrieb Götz Reinicke < goetz.reinicke@filmakademie.de>:
Hi,
As I’v read and thought a lot about the migration as this is a bigger project, I was wondering if anyone has done that already and might share some notes or playbooks, because in all readings there where some parts missing or miss understandable to me.
I do have some different approaches in mind, so may be you have some suggestions or hints.
a) upgrade nautilus on centos 7 with the few missing features like dashboard and prometheus. After that migrate one node after an other to ubuntu 20.04 with octopus and than upgrade ceph to the recent stable version.
b) migrate one node after an other to ubuntu 18.04 with nautilus and then upgrade to octupus and after that to ubuntu 20.04.
or
c) upgrade one node after an other to ubuntu 20.04 with octopus and join it to the cluster until all nodes are upgraded.
For test I tried c) with a mon node, but adding that to the cluster fails with some failed state, still probing for the other mons. (I dont have the right log at hand right now.)
So my questions are:
a) What would be the best (most stable) migration path and
b) is it in general possible to add a new octopus mon (not upgraded one) to a nautilus cluster, where the other mons are still on nautilus?
I hope my thoughts and questions are understandable :)
Thanks for any hint and suggestion. Best . Götz _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi, We went the „long“ way. - first emptied osd node by node (for each pool), purged all OSDs - moved the OS from centos 7 to ubuntu 20 (reinstalled every node) - removed the cache pool and cleaned up some config - installed all OSDs and moved the data back - upgraded ceph nautilus to octopus (containered) and now we are moving to pacific/quincy. This worked (slow - lot’s of data) without any bigger outage or problems. The only thing was, that a pool of NVMe nodes was flooding other nodes while backfilling :) … … so start with low backfill values and CHECK before you carrie on with other pools or nodes. Best . Götz
Am 17.01.2024 um 08:09 schrieb Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com>:
Hi Goetz,
Which method you finally choose? We've done a successful migration from Centos 8 to ubuntu 20.04 but we have a centos 7 nautilus cluster which we'd like to move to Ubuntu 20.04 octopus same as you. Wonder any of you tried to skip Rocky 8 from the flow?
<…>
As I’v read and thought a lot about the migration as this is a bigger project, I was wondering if anyone has done that already and might share some notes or playbooks, because in all readings there where some parts missing or miss understandable to me.
I do have some different approaches in mind, so may be you have some suggestions or hints.
I think I will update the el7 nodes to el9 with nautilus. It seems possible to build the nautilus rpms for el9. After that upgrading ceph. This will prevent having to upgrade the nodes 2 times.
Hi Götz, We’ve done a similar process which involves going from starting at CentOS 7 Nautilus and upgrading to Rocky 8/Ubuntu 20.04 Octopus+. What we do is start on CentOS 7 Nautilus we upgrade to Octopus on CentOS 7 (we’ve built python packages and have them on our repo to satisfy some ceph-mgr things and such with octopus on centos 7) From here we have a process to migrate the node from CentOS 7 to Rocky 8 preserving the OS and stuff, but you could also just reinstall the OS and reinstall ceph packages/config files if need be. Once on Rocky 8 and Octopus we then upgrade the ceph versions further. The order of the upgrades would be like this: CentOS 7 / Nautilus CentOS 7 / Octopus Rocky 8 / Octopus Ubuntu 20 / Octopus We aren’t changing from CentOS right to Ubuntu OS but process is similar enough. The one time we did switch to Ubuntu we just did the same process then once on the latest ceph version just reinstalled a node at a time from Rocky to Ubuntu. Probably are fine to go right from CentOS 7 to Ubuntu but we figured it’d be more reliable to go from el8 rpms to debs than el7 rpms to debs. I would say make the priority matching ceph versions if possible, and then OS. This is what has worked for us. Other people may have different experiences however. In your case, your b choice is the closest to what we would do so I would say that should be the safest. Overall your a and b is kind of a mix of what we do for our upgrades. Regards, Bailey
Hi, from Ceph perspective it's supported to upgrade from N to P, you can safely skip O. We have done that on several clusters without any issues. You just need to make sure that your upgrade to N was complete. Just a few days ago someone tried to upgrade from O to Q with "require-osd-release nautilus" which broke the upgrade. We have plans to switch our OS as well (probably on some customer clusters as well) but since they all are already under cephadm management I don't expect major issues with that. Regards, Eugen Zitat von Bailey Allison <ballison@45drives.com>:
Hi Götz,
We’ve done a similar process which involves going from starting at CentOS 7 Nautilus and upgrading to Rocky 8/Ubuntu 20.04 Octopus+.
What we do is start on CentOS 7 Nautilus we upgrade to Octopus on CentOS 7 (we’ve built python packages and have them on our repo to satisfy some ceph-mgr things and such with octopus on centos 7)
From here we have a process to migrate the node from CentOS 7 to Rocky 8 preserving the OS and stuff, but you could also just reinstall the OS and reinstall ceph packages/config files if need be.
Once on Rocky 8 and Octopus we then upgrade the ceph versions further.
The order of the upgrades would be like this:
CentOS 7 / Nautilus
CentOS 7 / Octopus
Rocky 8 / Octopus
Ubuntu 20 / Octopus
We aren’t changing from CentOS right to Ubuntu OS but process is similar enough.
The one time we did switch to Ubuntu we just did the same process then once on the latest ceph version just reinstalled a node at a time from Rocky to Ubuntu. Probably are fine to go right from CentOS 7 to Ubuntu but we figured it’d be more reliable to go from el8 rpms to debs than el7 rpms to debs.
I would say make the priority matching ceph versions if possible, and then OS. This is what has worked for us. Other people may have different experiences however.
In your case, your b choice is the closest to what we would do so I would say that should be the safest.
Overall your a and b is kind of a mix of what we do for our upgrades.
Regards,
Bailey
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
It's all covered in the docs [1], one of the points I already mentioned (require-osd-release), you should have bluestore OSDs and converted them to ceph-volume before you can adopt them with cephadm (if you deployed your cluster pre-nautilus). [1] https://docs.ceph.com/en/nautilus/releases/nautilus/#upgrading-from-mimic-or... Zitat von Marc <Marc@f1-outsourcing.eu>:
from Ceph perspective it's supported to upgrade from N to P, you can safely skip O. We have done that on several clusters without any issues. You just need to make sure that your upgrade to N was complete.
How do you verify if the upgrade was complete?
We went through this exercise, though our starting point was ubuntu 16.04 / nautilus. We reduced our double builds as follows: 1. Rebuild each monitor host on 18.04/bionic and rejoin still on nautilus 2. Upgrade all mons, mgrs., (and rgws optionally) to pacific 3. Convert each mon, mgr, rgw to cephadm and enable orchestrator 4. Rebuild each mon, mgr, rgw on 20.04/focal and rejoin pacfic cluster 5. Drain and rebuild each osd host on focal and pacific This has the advantage of only having to drain and rebuild the OSD hosts once. Double building the control cluster hosts isn’t so bad, and orchestrator makes all of the ceph parts easy once it’s enabled. The biggest challenge we ran into was: https://tracker.ceph.com/issues/51652 because we still had a lot of filestore osds. It’s frustrating, but we managed to get through it without much client interruption on a dozen prod clusters, most of which were 38 osd hosts and 912 total osds each. One thing which helped, was, before beginning the osd host builds, set all of the old osds primary-affinity to something <1. This way when the new pacific (or octopus) osds join the cluster they will automatically be favored for primary on their pgs. If a heartbeat timeout storm starts to get out of control, start by setting nodown and noout. The flapping osds are the worst. Then figure out which osds are the culprit and restart them. Hopefully your nautilus osds are all bluestore and you won’t have this problem. We put up with it, because the filestore to bluestore conversion was one of the most important parts of this upgrade for us. Best of luck, whatever route you take. Regards, Josh Beaman From: Götz Reinicke <goetz.reinicke@filmakademie.de> Date: Tuesday, August 1, 2023 at 1:01 PM To: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [EXTERNAL] [ceph-users] Upgrading nautilus / centos7 to octopus / ubuntu 20.04. - Suggestions and hints? Hi, As I’v read and thought a lot about the migration as this is a bigger project, I was wondering if anyone has done that already and might share some notes or playbooks, because in all readings there where some parts missing or miss understandable to me. I do have some different approaches in mind, so may be you have some suggestions or hints. a) upgrade nautilus on centos 7 with the few missing features like dashboard and prometheus. After that migrate one node after an other to ubuntu 20.04 with octopus and than upgrade ceph to the recent stable version. b) migrate one node after an other to ubuntu 18.04 with nautilus and then upgrade to octupus and after that to ubuntu 20.04. or c) upgrade one node after an other to ubuntu 20.04 with octopus and join it to the cluster until all nodes are upgraded. For test I tried c) with a mon node, but adding that to the cluster fails with some failed state, still probing for the other mons. (I dont have the right log at hand right now.) So my questions are: a) What would be the best (most stable) migration path and b) is it in general possible to add a new octopus mon (not upgraded one) to a nautilus cluster, where the other mons are still on nautilus? I hope my thoughts and questions are understandable :) Thanks for any hint and suggestion. Best . Götz
Hi, thanks to all suggestions. Right now, it is step by step that works: going to bionic/nautilus …and from that like Josh noted. We encountered a problem which I'll post separately . Best . Götz
Am 03.08.2023 um 15:44 schrieb Beaman, Joshua <Joshua_Beaman@comcast.com>:
We went through this exercise, though our starting point was ubuntu 16.04 / nautilus. We reduced our double builds as follows:
Rebuild each monitor host on 18.04/bionic and rejoin still on nautilus Upgrade all mons, mgrs., (and rgws optionally) to pacific Convert each mon, mgr, rgw to cephadm and enable orchestrator Rebuild each mon, mgr, rgw on 20.04/focal and rejoin pacfic cluster Drain and rebuild each osd host on focal and pacific
This has the advantage of only having to drain and rebuild the OSD hosts once. Double building the control cluster hosts isn’t so bad, and orchestrator makes all of the ceph parts easy once it’s enabled.
The biggest challenge we ran into was: https://tracker.ceph.com/issues/51652 because we still had a lot of filestore osds. It’s frustrating, but we managed to get through it without much client interruption on a dozen prod clusters, most of which were 38 osd hosts and 912 total osds each. One thing which helped, was, before beginning the osd host builds, set all of the old osds primary-affinity to something <1. This way when the new pacific (or octopus) osds join the cluster they will automatically be favored for primary on their pgs. If a heartbeat timeout storm starts to get out of control, start by setting nodown and noout. The flapping osds are the worst. Then figure out which osds are the culprit and restart them.
Hopefully your nautilus osds are all bluestore and you won’t have this problem. We put up with it, because the filestore to bluestore conversion was one of the most important parts of this upgrade for us.
Best of luck, whatever route you take.
Regards, Josh Beaman
participants (7)
-
Bailey Allison
-
Beaman, Joshua
-
Boris Behrens
-
Eugen Block
-
Götz Reinicke
-
Marc
-
Szabo, Istvan (Agoda)