Hello, I'm curious how the community handles backup and disaster recovery planning with Ceph in production environments. Does running Ceph fundamentally change your backup strategy compared to traditional storage solutions? Specifically, I'm interested in: - Are you backing up at the RBD level, or relying on application-level backups? - How do you handle cross-site replication for disaster recovery? - Any tools or practices you'd recommend for automated backup scheduling? Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai <mailto:anthony@linehaul.ai> [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181 [image] www.linehaul.ai <http://www.linehaul.ai/> [image] <http://www.linehaul.ai/> [image] <https://www.linkedin.com/in/anthony-fec/>
Hi, Anthony Fecarotta schreef op 2025-08-13 18:36:
Hello,
I'm curious how the community handles backup and disaster recovery planning with Ceph in production environments. Does running Ceph fundamentally change your backup strategy compared to traditional storage solutions?
Specifically, I'm interested in: - Are you backing up at the RBD level, or relying on application-level backups?
I would say both are critical. Especially when doing application-level clustering; you may not be able to get away with restoring an entire VM.
- How do you handle cross-site replication for disaster recovery?
Are you talking about running Ceph in a multisite fashion, or just storing backups / RBD images somewhere else?
- Any tools or practices you'd recommend for automated backup scheduling?
When using Proxmox with Ceph, Proxmox Backup Server (PBS) is a no-brainer. I know that there's enterprise-y solutions like Storware. I'm not sure how non-enterprise-minded folks approach this, though.
Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai <mailto:anthony@linehaul.ai> [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181
[image] www.linehaul.ai <http://www.linehaul.ai/> [image] <http://www.linehaul.ai/> [image] <https://www.linkedin.com/in/anthony-fec/> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Met vriendelijke groeten, William David Edwards
Are you talking about running Ceph in a multisite fashion, or just storing backups / RBD images somewhere else? Multisite.
When using Proxmox with Ceph, Proxmox Backup Server (PBS) is a no-brainer. Currently using PBS on baremetal, and I don't plan on stopping, as long as Proxmox is part of our stack; but I do would like to either: A. replicate that PBS server to another server off-site. B. Backup directly to secondary, off-site backup; but it would be a separate storage solution.
Option B may seem non-sensical, but something about diversifying my backup software makes me feel better — or maybe I just have too much fun using different platforms. I have heard good thing about Storware, but I won't put any data on a public cloud, if I can help it. I've been through corporate lawsuits, and I've seen entire databases get turned over to discovery when only a few rows were even requested by opposing counsel. Then again, there is encryption.... for now. [image] Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai <mailto:anthony@linehaul.ai> [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181 [image] www.linehaul.ai <http://www.linehaul.ai/> [image] <http://www.linehaul.ai/> [image] <https://www.linkedin.com/in/anthony-fec/> On Wed Aug 13, 2025, 05:26 PM GMT, William David Edwards <mailto:wedwards@cyberfusion.nl> wrote:
Hi,
Anthony Fecarotta schreef op 2025-08-13 18:36:
Hello,
I'm curious how the community handles backup and disaster recovery planning with Ceph in production environments. Does running Ceph fundamentally change your backup strategy compared to traditional storage solutions?
Specifically, I'm interested in: - Are you backing up at the RBD level, or relying on application-level backups?
I would say both are critical. Especially when doing application-level clustering; you may not be able to get away with restoring an entire VM.
- How do you handle cross-site replication for disaster recovery?
Are you talking about running Ceph in a multisite fashion, or just storing backups / RBD images somewhere else?
- Any tools or practices you'd recommend for automated backup scheduling?
When using Proxmox with Ceph, Proxmox Backup Server (PBS) is a no-brainer.
I know that there's enterprise-y solutions like Storware. I'm not sure how non-enterprise-minded folks approach this, though.
Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai <mailto:anthony@linehaul.ai> [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181
[image] www.linehaul.ai <http://www.linehaul.ai/> [image] <http://www.linehaul.ai/> [image] <https://www.linkedin.com/in/anthony-fec/> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Met vriendelijke groeten,
William David Edwards
I'm using Bacula on the Ceph filesystem. Not presently backing up object store since I really only experiment with it, but Bacula can handle that. Bacula is probably not the most polished system I could use, but unlike some commercial ones I've used, it has always delivered and has the virtue that in case of Total Nuclear Destruction I can still restore backups using offline utilities. And is part of the standard package set for Red Hat family distros. Because Ceph has proven so reliable, I really have only needed it to recover goofed-up files rather than from host failure. a local git archive helps in really critical stuff, too. My systems back up incrementally (ceph and local filesystems) at around 4am local time daily with a full backup once a week. First-stage backup is to a local server local disk, then offline copies go out to other media. I'd love to have a robotic tape library like I used to have, but current economics make that solution impossible. Now if I could just archive to holographic diamond storage... Tim On 8/13/25 12:36, Anthony Fecarotta wrote:
Hello,
I'm curious how the community handles backup and disaster recovery planning with Ceph in production environments. Does running Ceph fundamentally change your backup strategy compared to traditional storage solutions?
Specifically, I'm interested in: - Are you backing up at the RBD level, or relying on application-level backups? - How do you handle cross-site replication for disaster recovery? - Any tools or practices you'd recommend for automated backup scheduling?
Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai <mailto:anthony@linehaul.ai> [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181
[image] www.linehaul.ai <http://www.linehaul.ai/> [image] <http://www.linehaul.ai/> [image] <https://www.linkedin.com/in/anthony-fec/> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I've been, totally independently, also using Bacula and doing instance filesystem backups, not objects. I also have only needed to restore file-level restores occasionally as a result of end-user error. I use snapshots before major upgrades or migrations but these are all manual. The only other difference is I copy bacula volumes are stored off-site in AWS Glacier with a 6 month retention. It checks the boxes for DR and gets tested semi-regularly by end-user errors. peter On Wed, Aug 13, 2025 at 02:03 PM, Tim Holloway wrote: I'm using Bacula on the Ceph filesystem. Not presently backing up object store since I really only experiment with it, but Bacula can handle that. Bacula is probably not the most polished system I could use, but unlike some commercial ones I've used, it has always delivered and has the virtue that in case of Total Nuclear Destruction I can still restore backups using offline utilities. And is part of the standard package set for Red Hat family distros. Because Ceph has proven so reliable, I really have only needed it to recover goofed-up files rather than from host failure. a local git archive helps in really critical stuff, too. My systems back up incrementally (ceph and local filesystems) at around 4am local time daily with a full backup once a week. First-stage backup is to a local server local disk, then offline copies go out to other media. I'd love to have a robotic tape library like I used to have, but current economics make that solution impossible. Now if I could just archive to holographic diamond storage... Tim On 8/13/25 12:36, Anthony Fecarotta wrote: Hello, I'm curious how the community handles backup and disaster recovery planning with Ceph in production environments. Does running Ceph fundamentally change your backup strategy compared to traditional storage solutions? Specifically, I'm interested in: - Are you backing up at the RBD level, or relying on application-level backups? - How do you handle cross-site replication for disaster recovery? - Any tools or practices you'd recommend for automated backup scheduling? Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai (mailto:anthony@linehaul.ai) [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181 [image] www.linehaul.ai (http://www.linehaul.ai) [image] [image] _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io (mailto:ceph-users@ceph.io) To unsubscribe send an email to ceph-users-leave@ceph.io (mailto:ceph-users-leave@ceph.io) _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io (mailto:ceph-users@ceph.io) To unsubscribe send an email to ceph-users-leave@ceph.io (mailto:ceph-users-leave@ceph.io)
I modified for my purposes this little tool: https://github.com/teralytics/ceph-backup to do incremental backup for RBD volumes. In addition I wrote also a little python tool to flatten rbd incremental snapshot downloads over the corresponding full rbd image (My backup appliance is mounted as a filesystem). With this strategy I do a single full rbd image download and subsequently I need only to do incremental snapshot downloads; only a predefined bunch of snapshots are stored along the full image by flattening the oldest one over the full-image, then the corresponding snapshot is removed (daily backup). As we use ceph for openstack volumes, we did also a very little change into cinder code, and all of the backup works are completely invisible to people, just admin knows that there are the backup-snapshots. Probably rough, but works: If anyone is interested in, just ask. My 5cents... On 08/13/2025 06:36 PM, Anthony Fecarotta wrote:
Hello,
I'm curious how the community handles backup and disaster recovery planning with Ceph in production environments. Does running Ceph fundamentally change your backup strategy compared to traditional storage solutions?
Specifically, I'm interested in: - Are you backing up at the RBD level, or relying on application-level backups? - How do you handle cross-site replication for disaster recovery? - Any tools or practices you'd recommend for automated backup scheduling?
Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai <mailto:anthony@linehaul.ai> [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181
[image] www.linehaul.ai <http://www.linehaul.ai/> [image] <http://www.linehaul.ai/> [image] <https://www.linkedin.com/in/anthony-fec/> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- ing. Sergio Rabellino Università degli Studi di Torino Dipartimento di Informatica Tecnico di Ricerca - Servizio ICT Cell +342-529-5409 Tel +39-0116706701 Fax +39-011751603 C.so Svizzera , 185 - 10149 - Torino <http://www.di.unito.it>
I love the ingenuity. I will most likely be reaching out at some point in the future, if not just out of sheer curiosity. Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai <mailto:anthony@linehaul.ai> [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181 [image] www.linehaul.ai <http://www.linehaul.ai/> [image] <http://www.linehaul.ai/> [image] <https://www.linkedin.com/in/anthony-fec/> On Wed Aug 13, 2025, 10:04 PM GMT, Sergio Rabellino <mailto:rabellino@di.unito.it> wrote:
I modified for my purposes this little tool: https://github.com/teralytics/ceph-backup to do incremental backup for RBD volumes.
In addition I wrote also a little python tool to flatten rbd incremental snapshot downloads over the corresponding full rbd image (My backup appliance is mounted as a filesystem).
With this strategy I do a single full rbd image download and subsequently I need only to do incremental snapshot downloads; only a predefined bunch of snapshots are stored along the full image by flattening the oldest one over the full-image, then the corresponding snapshot is removed (daily backup).
As we use ceph for openstack volumes, we did also a very little change into cinder code, and all of the backup works are completely invisible to people, just admin knows that there are the backup-snapshots.
Probably rough, but works: If anyone is interested in, just ask.
My 5cents...
On 08/13/2025 06:36 PM, Anthony Fecarotta wrote:
Hello,
I'm curious how the community handles backup and disaster recovery planning with Ceph in production environments. Does running Ceph fundamentally change your backup strategy compared to traditional storage solutions?
Specifically, I'm interested in: - Are you backing up at the RBD level, or relying on application-level backups? - How do you handle cross-site replication for disaster recovery? - Any tools or practices you'd recommend for automated backup scheduling?
Regards, [image] Anthony Fecarotta Founder & President [image] anthony@linehaul.ai <mailto:anthony@linehaul.ai> [image] 224-339-1182 [image] (855) 625-0300 [image] 1 Mid America Plz Flr 3 Oakbrook Terrace, IL 60181
[image] www.linehaul.ai <http://www.linehaul.ai/> [image] <http://www.linehaul.ai/> [image] <https://www.linkedin.com/in/anthony-fec/> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- ing. Sergio Rabellino
Università degli Studi di Torino Dipartimento di Informatica Tecnico di Ricerca - Servizio ICT Cell +342-529-5409 Tel +39-0116706701 Fax +39-011751603 C.so Svizzera , 185 - 10149 - Torino
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (5)
-
Anthony Fecarotta
-
Peter Eisch
-
Sergio Rabellino
-
Tim Holloway
-
William David Edwards