Can Ceph Do The Job?
We are looking to role out a all flash Ceph cluster as storage for our cloud solution. The OSD's will be on slightly slower Micron 5300 PRO's, with WAL/DB on Micron 7300 MAX NVMe's. My main concern with Ceph being able to fit the bill is its snapshot abilities. For each RBD we would like the following snapshots 8x 30 minute snapshots (latest 4 hours) With our current solution (HPE Nimble) we simply pause all write IO on the 10 minute mark for roughly 2 seconds and then we take a snapshot of the entire Nimble volume. Each VM within the Nimble volume is sitting on a Linux Logical Volume so its easy for us to take one big snapshot and only get access to a specific clients data. Are there any options for automating managing/retention of snapshots within Ceph besides some bash scripts? Is there anyway to take snapshots of all RBD's within a pool at a given time? Is there anyone successfully running with this many snapshots? If anyone is running a similar setup, would love to hear how your doing it.
We are making hourly snapshots of ~400 rbd drives in one (spinning-rust) cluster. The snapshots are made one by one. Total size of the base images is around 80TB. The entire process takes a few minutes. We do not experience any problems doing this. Op do 30 jan. 2020 om 15:30 schreef Adam Boyhan <adamb@medent.com>:
We are looking to role out a all flash Ceph cluster as storage for our cloud solution. The OSD's will be on slightly slower Micron 5300 PRO's, with WAL/DB on Micron 7300 MAX NVMe's.
My main concern with Ceph being able to fit the bill is its snapshot abilities.
For each RBD we would like the following snapshots
8x 30 minute snapshots (latest 4 hours)
With our current solution (HPE Nimble) we simply pause all write IO on the 10 minute mark for roughly 2 seconds and then we take a snapshot of the entire Nimble volume. Each VM within the Nimble volume is sitting on a Linux Logical Volume so its easy for us to take one big snapshot and only get access to a specific clients data.
Are there any options for automating managing/retention of snapshots within Ceph besides some bash scripts? Is there anyway to take snapshots of all RBD's within a pool at a given time?
Is there anyone successfully running with this many snapshots? If anyone is running a similar setup, would love to hear how your doing it. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Bastiaan Visser (bastiaan) writes:
We are making hourly snapshots of ~400 rbd drives in one (spinning-rust) cluster. The snapshots are made one by one. Total size of the base images is around 80TB. The entire process takes a few minutes. We do not experience any problems doing this.
Out of curiosity - which version of CEPH ? I've read that I/O to RBD images for which snapshots exist is impacted. Is that your experience ? And, do you maintain the snapshots, or are they discarded after (i.e.: are they used as potential point-in-time recovery points, or are they used to take a backup, then discarded ?) Cheers, Phil
Den tors 30 jan. 2020 kl 15:29 skrev Adam Boyhan <adamb@medent.com>:
We are looking to role out a all flash Ceph cluster as storage for our cloud solution. The OSD's will be on slightly slower Micron 5300 PRO's, with WAL/DB on Micron 7300 MAX NVMe's. My main concern with Ceph being able to fit the bill is its snapshot abilities. For each RBD we would like the following snapshots 8x 30 minute snapshots (latest 4 hours) With our current solution (HPE Nimble) we simply pause all write IO on the 10 minute mark for roughly 2 seconds and then we take a snapshot of the entire Nimble volume. Each VM within the Nimble volume is sitting on a Linux Logical Volume so its easy for us to take one big snapshot and only get access to a specific clients data. Are there any options for automating managing/retention of snapshots within Ceph besides some bash scripts? Is there anyway to take snapshots of all RBD's within a pool at a given time?
You could make a snapshot of the whole pool, that would cover all RBDs in it I gather? https://docs.ceph.com/docs/nautilus/rados/operations/pools/#make-a-snapshot-... But if you need to work in parallel with each snapshot from different times and clone them one by one and so forth, doing it per-RBD would be better. https://docs.ceph.com/docs/nautilus/rbd/rbd-snapshot/ -- May the most significant bit of your life be positive.
Its my understanding that pool snapshots would basically require us to be in a all or nothing situation were we would have to revert all RBD's in a pool. If we could clone a pool snapshot for filesystem level access like a rbd snapshot, that would help a ton. Thanks, Adam Boyhan System & Network Administrator MEDENT(EMR/EHR) 15 Hulbert Street - P.O. Box 980 Auburn, New York 13021 www.medent.com Phone: (315)-255-1751 Fax: (315)-255-3539 Cell: (315)-729-2290 adamb@medent.com This message and any attachments may contain information that is protected by law as privileged and confidential, and is transmitted for the sole use of the intended recipient(s). If you are not the intended recipient, you are hereby notified that any use, dissemination, copying or retention of this e-mail or the information contained herein is strictly prohibited. If you received this e-mail in error, please immediately notify the sender by e-mail, and permanently delete this e-mail. From: "Janne Johansson" <icepic.dz@gmail.com> To: "adamb" <adamb@medent.com> Cc: "ceph-users" <ceph-users@ceph.io> Sent: Thursday, January 30, 2020 10:06:14 AM Subject: Re: [ceph-users] Can Ceph Do The Job? Den tors 30 jan. 2020 kl 15:29 skrev Adam Boyhan < [ mailto:adamb@medent.com | adamb@medent.com ] >: We are looking to role out a all flash Ceph cluster as storage for our cloud solution. The OSD's will be on slightly slower Micron 5300 PRO's, with WAL/DB on Micron 7300 MAX NVMe's. My main concern with Ceph being able to fit the bill is its snapshot abilities. For each RBD we would like the following snapshots 8x 30 minute snapshots (latest 4 hours) With our current solution (HPE Nimble) we simply pause all write IO on the 10 minute mark for roughly 2 seconds and then we take a snapshot of the entire Nimble volume. Each VM within the Nimble volume is sitting on a Linux Logical Volume so its easy for us to take one big snapshot and only get access to a specific clients data. Are there any options for automating managing/retention of snapshots within Ceph besides some bash scripts? Is there anyway to take snapshots of all RBD's within a pool at a given time? You could make a snapshot of the whole pool, that would cover all RBDs in it I gather? [ https://docs.ceph.com/docs/nautilus/rados/operations/pools/#make-a-snapshot-... | https://docs.ceph.com/docs/nautilus/rados/operations/pools/#make-a-snapshot-... ] But if you need to work in parallel with each snapshot from different times and clone them one by one and so forth, doing it per-RBD would be better. [ https://docs.ceph.com/docs/nautilus/rbd/rbd-snapshot/ | https://docs.ceph.com/docs/nautilus/rbd/rbd-snapshot/ ] -- May the most significant bit of your life be positive.
participants (4)
-
Adam Boyhan
-
Bastiaan Visser
-
Janne Johansson
-
Phil Regnauld