Patrick brought up a good point in today’s Steering Committee call about migrating the LRC. There is value in having a cluster that has gone through dozens of Ceph releases/upgrades. I see a few options: 1. Do not buy new LRC hardware for POK. Just use Proxmox storage for teuthology logs, etc. until Sepia LRC gear moves 2. Buy new LRC gear, migrate what’s necessary, integrate old Sepia gear into “new” LRC 3. Buy new LRC gear, have new modern LRC and run alongside old LRC There is no scenario where we don’t move the existing LRC in my opinion. Thoughts? -- David Galloway Ceph Engineering Labs – Infrastructure Architect +1 989 295 0091 - Mobile david.galloway@ibm.com IBM
On Mon, Mar 24, 2025 at 12:02 PM David Galloway <David.Galloway@ibm.com> wrote:
Buy new LRC gear, migrate what’s necessary, integrate old Sepia gear into “new” LRC
I don't think we necessarily need to keep the old sepia LRC nodes. As long as we migrate everything to new hardware it should be good. I think the LRC can just steal 3-6 nodes from the hardware we already plan to purchase (for the upstream lab). I think the troublesome part will be the RADOS replication to the new site. Will we have the VPN straddling the two sites at some point? -- Patrick Donnelly, Ph.D. He / Him / His Red Hat Partner Engineer IBM, Inc. GPG: 19F28A586F808C2402351B93C3301A3E258DD79D
If we want to keep the existing LRC (which would be great), the important part is the internal state of the system. So as Patrick alludes to, we *could* create a single cluster that is stretched across old and new labs, and then use CRUSH rule changes to move everything into the new space. I'm not sure this is easier than simply shipping the existing nodes intact and then using them to seed a new LRC in-place, if we have a solution to store what we need in the lab until that shipping happens. To keep the historical aging of the cluster, we do need the existing LRC nodes to be part of it when we generate the cluster, though — we can't merge them later in any meaningful way. I'm not sure if that changes your perception of how reasonable those options are, David. -Greg On Mon, Mar 24, 2025 at 9:17 AM Patrick Donnelly <pdonnell@redhat.com> wrote:
On Mon, Mar 24, 2025 at 12:02 PM David Galloway <David.Galloway@ibm.com> wrote:
Buy new LRC gear, migrate what’s necessary, integrate old Sepia gear into “new” LRC
I don't think we necessarily need to keep the old sepia LRC nodes. As long as we migrate everything to new hardware it should be good. I think the LRC can just steal 3-6 nodes from the hardware we already plan to purchase (for the upstream lab).
I think the troublesome part will be the RADOS replication to the new site. Will we have the VPN straddling the two sites at some point?
-- Patrick Donnelly, Ph.D. He / Him / His Red Hat Partner Engineer IBM, Inc. GPG: 19F28A586F808C2402351B93C3301A3E258DD79D _______________________________________________ Sepia mailing list -- sepia@ceph.io To unsubscribe send an email to sepia-leave@ceph.io
We are able to bring the Sepia switches with us so we are not limited on network ports, space, power, or cooling. We only have a 1Gb link into the Community Cage and we share with the other tenants so replication will be limited there. If we end up using Proxmox, the backing storage will be Ceph anyway so we will have a Ceph cluster in Poughkeepsie before we ship the current Sepia gear. Good to know that we can’t add the nodes later. I’m thinking let’s (eventually) shrink the cluster to be a minimum number of useful hosts (all ivan + N reesi) and we can chuck them all in a rack together. Anything running (e.g., teuth logs, home dirs) on the “Proxmox” Ceph cluster can be copied over after. From: Gregory Farnum <gfarnum@redhat.com> Date: Tuesday, March 25, 2025 at 1:42 AM To: Patrick Donnelly <pdonnell@redhat.com> Cc: David Galloway <David.Galloway@ibm.com>, sepia@ceph.com <sepia@ceph.io> Subject: [EXTERNAL] Re: [sepia] Re: Migrating the LRC to IBM's lab If we want to keep the existing LRC (which would be great), the important part is the internal state of the system. So as Patrick alludes to, we *could* create a single cluster that is stretched across old and new labs, and then use CRUSH rule changes to move everything into the new space. I'm not sure this is easier than simply shipping the existing nodes intact and then using them to seed a new LRC in-place, if we have a solution to store what we need in the lab until that shipping happens. To keep the historical aging of the cluster, we do need the existing LRC nodes to be part of it when we generate the cluster, though — we can't merge them later in any meaningful way. I'm not sure if that changes your perception of how reasonable those options are, David. -Greg On Mon, Mar 24, 2025 at 9:17 AM Patrick Donnelly <pdonnell@redhat.com> wrote:
On Mon, Mar 24, 2025 at 12:02 PM David Galloway <David.Galloway@ibm.com> wrote:
Buy new LRC gear, migrate what’s necessary, integrate old Sepia gear into “new” LRC
I don't think we necessarily need to keep the old sepia LRC nodes. As long as we migrate everything to new hardware it should be good. I think the LRC can just steal 3-6 nodes from the hardware we already plan to purchase (for the upstream lab).
I think the troublesome part will be the RADOS replication to the new site. Will we have the VPN straddling the two sites at some point?
-- Patrick Donnelly, Ph.D. He / Him / His Red Hat Partner Engineer IBM, Inc. GPG: 19F28A586F808C2402351B93C3301A3E258DD79D _______________________________________________ Sepia mailing list -- sepia@ceph.io To unsubscribe send an email to sepia-leave@ceph.io
On Tue, Mar 25, 2025 at 11:31 AM David Galloway <David.Galloway@ibm.com> wrote:
We are able to bring the Sepia switches with us so we are not limited on network ports, space, power, or cooling.
We only have a 1Gb link into the Community Cage and we share with the other tenants so replication will be limited there.
If we end up using Proxmox, the backing storage will be Ceph anyway so we will have a Ceph cluster in Poughkeepsie before we ship the current Sepia gear.
Good to know that we can’t add the nodes later. I’m thinking let’s (eventually) shrink the cluster to be a minimum number of useful hosts (all ivan + N reesi) and we can chuck them all in a rack together. Anything running (e.g., teuth logs, home dirs) on the “Proxmox” Ceph cluster can be copied over after.
That sounds like a good plan to me!
*From: *Gregory Farnum <gfarnum@redhat.com> *Date: *Tuesday, March 25, 2025 at 1:42 AM *To: *Patrick Donnelly <pdonnell@redhat.com> *Cc: *David Galloway <David.Galloway@ibm.com>, sepia@ceph.com < sepia@ceph.io> *Subject: *[EXTERNAL] Re: [sepia] Re: Migrating the LRC to IBM's lab
If we want to keep the existing LRC (which would be great), the important part is the internal state of the system. So as Patrick alludes to, we *could* create a single cluster that is stretched across old and new labs, and then use CRUSH rule changes to move everything into the new space.
I'm not sure this is easier than simply shipping the existing nodes intact and then using them to seed a new LRC in-place, if we have a solution to store what we need in the lab until that shipping happens.
To keep the historical aging of the cluster, we do need the existing LRC nodes to be part of it when we generate the cluster, though — we can't merge them later in any meaningful way. I'm not sure if that changes your perception of how reasonable those options are, David. -Greg
On Mon, Mar 24, 2025 at 9:17 AM Patrick Donnelly <pdonnell@redhat.com> wrote:
On Mon, Mar 24, 2025 at 12:02 PM David Galloway <David.Galloway@ibm.com>
wrote:
Buy new LRC gear, migrate what’s necessary, integrate old Sepia gear into “new” LRC
I don't think we necessarily need to keep the old sepia LRC nodes. As long as we migrate everything to new hardware it should be good. I think the LRC can just steal 3-6 nodes from the hardware we already plan to purchase (for the upstream lab).
I think the troublesome part will be the RADOS replication to the new site. Will we have the VPN straddling the two sites at some point?
-- Patrick Donnelly, Ph.D. He / Him / His Red Hat Partner Engineer IBM, Inc. GPG: 19F28A586F808C2402351B93C3301A3E258DD79D _______________________________________________ Sepia mailing list -- sepia@ceph.io To unsubscribe send an email to sepia-leave@ceph.io
Generally agree. LRC could also run in the cloud - I think it's value is less in which specific hardware it runs on and more on all the other aspects. On Mon, Mar 24, 2025 at 9:02 AM David Galloway <David.Galloway@ibm.com> wrote:
Patrick brought up a good point in today’s Steering Committee call about migrating the LRC. There is value in having a cluster that has gone through dozens of Ceph releases/upgrades.
I see a few options:
1. Do *not *buy new LRC hardware for POK. Just use Proxmox storage for teuthology logs, etc. until Sepia LRC gear moves 2. Buy new LRC gear, migrate what’s necessary, integrate old Sepia gear into “new” LRC 3. Buy new LRC gear, have new modern LRC and run alongside old LRC
There is no scenario where we don’t move the existing LRC in my opinion.
Thoughts?
--
David Galloway
Ceph Engineering Labs – Infrastructure Architect
+1 989 295 0091 - Mobile
david.galloway@ibm.com
IBM _______________________________________________ Sepia mailing list -- sepia@ceph.io To unsubscribe send an email to sepia-leave@ceph.io
participants (4)
-
Brett Niver
-
David Galloway
-
Gregory Farnum
-
Patrick Donnelly