ceph-volume simple disk scenario without LVM for OSD on PVC
Hi, I've started working on a saner way to deploy OSD with Rook so that they don't use the rook binary image. Why were/are we using the rook binary to activate the OSD? A bit of background on containers first, when executing a container, we need to provide a command entrypoint that will act as PID 1. So if you want to do pre/post action before running the process you need to use a wrapper. In Rook, that's the rook binary, which has a CLI and can then "activate" an OSD. Currently, this "rook osd activate" call does the following: * sed the lvm.conf * run c-v lvm activate * run the osd process On shutdown, we intercept the signal, "kill -9" the osd and de-activate the LV. I have a patch here: https://github.com/rook/rook/pull/4386, that solves the initial bullet points but one thing we cannot do is the signal catching and the lv de-activation. Before you ask, Kubernetes has pre/post-hook but they are not reliable, it's known and documented that there is no guarantee they would actually run before or after the container starts/stops. We tried and we had issues. Why do we want to stop using the rook binary for activation? Because each time we get a new binary version (new operator version), this will restart all the OSDs, even if the deployment spec didn't change, at least if nothing else than the rook image version changed. Also with containers, we have seen so many issues working with LVM, just to name a few: * adapt lvm filters * interactions with udev - need to tune the lvm config, even c-v itself has lvm flag to not sync with udev built-in * several bindmounts * lvm package must be present on the host even if running in containers * SELinux, yes lvm calls SELinux commands under the hood and pollute the logs in some scenarios Currently, one of the ways I can see this working is by not using LVM when bootstrapping OSDs. Unfortunately, some of the logic cannot go in the OSD code since the lv de-activation happens after the OSD stops. We need to de-activate the LV so when running in the Cloud the block can safely be re-attached to a new machine without LVM issues. I know this will be a bit challenging and might ultimately look like ceph-disk but it'd be nice to consider it. What about a small prototype for Bluestore with block/db/wal on the same disk? If this gets rejected, I might try a prototype for not using c-v in Rook or something else that might come up with this discussion. Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood."
On Tue, Dec 03, 2019 at 05:55:25PM +0100, Sebastien Han wrote:
Hi,
I've started working on a saner way to deploy OSD with Rook so that they don't use the rook binary image.
Why were/are we using the rook binary to activate the OSD?
A bit of background on containers first, when executing a container, we need to provide a command entrypoint that will act as PID 1. So if you want to do pre/post action before running the process you need to use a wrapper. In Rook, that's the rook binary, which has a CLI and can then "activate" an OSD. Currently, this "rook osd activate" call does the following:
* sed the lvm.conf * run c-v lvm activate * run the osd process
On shutdown, we intercept the signal, "kill -9" the osd and de-activate the LV.
I have a patch here: https://github.com/rook/rook/pull/4386, that solves the initial bullet points but one thing we cannot do is the signal catching and the lv de-activation. Before you ask, Kubernetes has pre/post-hook but they are not reliable, it's known and documented that there is no guarantee they would actually run before or after the container starts/stops. We tried and we had issues.
Why do we want to stop using the rook binary for activation? Because each time we get a new binary version (new operator version), this will restart all the OSDs, even if the deployment spec didn't change, at least if nothing else than the rook image version changed.
Also with containers, we have seen so many issues working with LVM, just to name a few:
* adapt lvm filters * interactions with udev - need to tune the lvm config, even c-v itself has lvm flag to not sync with udev built-in * several bindmounts * lvm package must be present on the host even if running in containers * SELinux, yes lvm calls SELinux commands under the hood and pollute the logs in some scenarios
I have only seen the last issue and that was a silly bug that was easily fixed. The others also sound like they can be fixed with reasonable effort. Is there anything that is technically hard to solve? It seems like dealing with config files and system infrastructure is just the normal pain of a deployment tool.
Currently, one of the ways I can see this working is by not using LVM when bootstrapping OSDs. Unfortunately, some of the logic cannot go in the OSD code since the lv de-activation happens after the OSD stops. We need to de-activate the LV so when running in the Cloud the block can safely be re-attached to a new machine without LVM issues.
I know this will be a bit challenging and might ultimately look like ceph-disk but it'd be nice to consider it. What about a small prototype for Bluestore with block/db/wal on the same disk?
If this gets rejected, I might try a prototype for not using c-v in Rook or something else that might come up with this discussion.
I have discussed this before (using bluestore) and I'm happy to write/look at a prototype. I don't however think that this solves all those issues listed once one factors in the feature set that user will expect. Just thinking about multi-device OSDs leads to partitions (which are much more fickle to setup) and encryption adds a ton of complexity. And those are features that user rely on today. Not to mention that we haven't even looked at some features yet (using dm-flakey for testing, caching on the block layer). Is the pain of using lvm so big that it seems unsolvable? I'd be happy to explore if this can't be solved with lvm in the picture.
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood." _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
-- Jan Fajerski Senior Software Engineer Enterprise Storage SUSE Software Solutions Germany GmbH Maxfeldstr. 5, 90409 Nürnberg, Germany (HRB 36809, AG Nürnberg) Geschäftsführer: Felix Imendörffer
@Alfredo, I haven't played with "simple" because the activation part is not an issue. Yes, the sub-command sounds like an option. @blaine, I know the logic in bash to intercept signals, we have something that does exactly all of that in ceph-container, it's just over-engineered in my opinion. And I don't want to rely on too much bash in Rook. @Sage, doing this in a new ceph-volume scenario is interesting, although ideally, I'd like to get rid of another wrapper... Also, this wrapper must be flexible and we must be able to change/adapt it quickly, that won't be the case if it lives in ceph-volume because it's in-tree Ceph (and I don't mean to resurrect another old discussion here ^^) @Jan, yes the biggest problem with the prototype is likely that once it works people will ask for more. I'm not saying lvm is unsolvable, nothing i. It's just that every time we have to work around it we add more and more complexity/requirements to the point where I'm having a hard time understand the value proposition. One thing, not ideal of course, would be to have a block implementation and lvm for the more advanced use cases because all the logic is present. I'm just afraid this might confuse (curious) users. Just to reiterate, I'm currently only looking at the simplest scenario which is the most common one: a Bluestore non-encrypted OSD with block/db/wal on the same disk. I'm going to investigate one more time why we need to de-activate the LV and see if I can find another fix so we don't have to do it explicitly. As for the signal catching, assuming I can fix the de-activation, this can go away with fast osd shutdown coming in Octopus. Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." On Tue, Dec 3, 2019 at 11:59 PM Jan Fajerski <jfajerski@suse.com> wrote:
On Tue, Dec 03, 2019 at 05:55:25PM +0100, Sebastien Han wrote:
Hi,
I've started working on a saner way to deploy OSD with Rook so that they don't use the rook binary image.
Why were/are we using the rook binary to activate the OSD?
A bit of background on containers first, when executing a container, we need to provide a command entrypoint that will act as PID 1. So if you want to do pre/post action before running the process you need to use a wrapper. In Rook, that's the rook binary, which has a CLI and can then "activate" an OSD. Currently, this "rook osd activate" call does the following:
* sed the lvm.conf * run c-v lvm activate * run the osd process
On shutdown, we intercept the signal, "kill -9" the osd and de-activate the LV.
I have a patch here: https://github.com/rook/rook/pull/4386, that solves the initial bullet points but one thing we cannot do is the signal catching and the lv de-activation. Before you ask, Kubernetes has pre/post-hook but they are not reliable, it's known and documented that there is no guarantee they would actually run before or after the container starts/stops. We tried and we had issues.
Why do we want to stop using the rook binary for activation? Because each time we get a new binary version (new operator version), this will restart all the OSDs, even if the deployment spec didn't change, at least if nothing else than the rook image version changed.
Also with containers, we have seen so many issues working with LVM, just to name a few:
* adapt lvm filters * interactions with udev - need to tune the lvm config, even c-v itself has lvm flag to not sync with udev built-in * several bindmounts * lvm package must be present on the host even if running in containers * SELinux, yes lvm calls SELinux commands under the hood and pollute the logs in some scenarios
I have only seen the last issue and that was a silly bug that was easily fixed. The others also sound like they can be fixed with reasonable effort. Is there anything that is technically hard to solve? It seems like dealing with config files and system infrastructure is just the normal pain of a deployment tool.
Currently, one of the ways I can see this working is by not using LVM when bootstrapping OSDs. Unfortunately, some of the logic cannot go in the OSD code since the lv de-activation happens after the OSD stops. We need to de-activate the LV so when running in the Cloud the block can safely be re-attached to a new machine without LVM issues.
I know this will be a bit challenging and might ultimately look like ceph-disk but it'd be nice to consider it. What about a small prototype for Bluestore with block/db/wal on the same disk?
If this gets rejected, I might try a prototype for not using c-v in Rook or something else that might come up with this discussion.
I have discussed this before (using bluestore) and I'm happy to write/look at a prototype. I don't however think that this solves all those issues listed once one factors in the feature set that user will expect. Just thinking about multi-device OSDs leads to partitions (which are much more fickle to setup) and encryption adds a ton of complexity. And those are features that user rely on today. Not to mention that we haven't even looked at some features yet (using dm-flakey for testing, caching on the block layer).
Is the pain of using lvm so big that it seems unsolvable? I'd be happy to explore if this can't be solved with lvm in the picture.
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood." _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
-- Jan Fajerski Senior Software Engineer Enterprise Storage SUSE Software Solutions Germany GmbH Maxfeldstr. 5, 90409 Nürnberg, Germany (HRB 36809, AG Nürnberg) Geschäftsführer: Felix Imendörffer
On 2019-12-04T10:11:42, Sebastien Han <shan@redhat.com> wrote:
Just to reiterate, I'm currently only looking at the simplest scenario which is the most common one: a Bluestore non-encrypted OSD with block/db/wal on the same disk.
This is not the most common scenario we see. While most are (so far) still not using at-rest encryption, the majority of our deployments has a separate WAL/DB (or journal, if from the old days). Regards, Lars -- SUSE Software Solutions Germany GmbH, MD: Felix Imendörffer, HRB 36809 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
We could play with making the entrypoing a simple bash script that traps the EXIT signal to kill the captured PID of the OSD. A `wait` might or might not be necessary. Similar thing in practice: http://mywiki.wooledge.org/SignalTrap#When_is_the_signal_handled.3F Blaine ________________________________ From: Sebastien Han <shan@redhat.com> Sent: Tuesday, December 3, 2019 09:55 To: dev@ceph.io <dev@ceph.io> Cc: Travis Nielsen <tnielsen@redhat.com> Subject: ceph-volume simple disk scenario without LVM for OSD on PVC Hi, I've started working on a saner way to deploy OSD with Rook so that they don't use the rook binary image. Why were/are we using the rook binary to activate the OSD? A bit of background on containers first, when executing a container, we need to provide a command entrypoint that will act as PID 1. So if you want to do pre/post action before running the process you need to use a wrapper. In Rook, that's the rook binary, which has a CLI and can then "activate" an OSD. Currently, this "rook osd activate" call does the following: * sed the lvm.conf * run c-v lvm activate * run the osd process On shutdown, we intercept the signal, "kill -9" the osd and de-activate the LV. I have a patch here: https://github.com/rook/rook/pull/4386, that solves the initial bullet points but one thing we cannot do is the signal catching and the lv de-activation. Before you ask, Kubernetes has pre/post-hook but they are not reliable, it's known and documented that there is no guarantee they would actually run before or after the container starts/stops. We tried and we had issues. Why do we want to stop using the rook binary for activation? Because each time we get a new binary version (new operator version), this will restart all the OSDs, even if the deployment spec didn't change, at least if nothing else than the rook image version changed. Also with containers, we have seen so many issues working with LVM, just to name a few: * adapt lvm filters * interactions with udev - need to tune the lvm config, even c-v itself has lvm flag to not sync with udev built-in * several bindmounts * lvm package must be present on the host even if running in containers * SELinux, yes lvm calls SELinux commands under the hood and pollute the logs in some scenarios Currently, one of the ways I can see this working is by not using LVM when bootstrapping OSDs. Unfortunately, some of the logic cannot go in the OSD code since the lv de-activation happens after the OSD stops. We need to de-activate the LV so when running in the Cloud the block can safely be re-attached to a new machine without LVM issues. I know this will be a bit challenging and might ultimately look like ceph-disk but it'd be nice to consider it. What about a small prototype for Bluestore with block/db/wal on the same disk? If this gets rejected, I might try a prototype for not using c-v in Rook or something else that might come up with this discussion. Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
It looks like preStop hooks might be useful if ceph daemons were to accept some sort of quit signal like nginx. https://chrislovecnm.com/kubernetes/best-practices/highly-available-resilien... ________________________________ From: Blaine Gardner <BlGardner@suse.com> Sent: Tuesday, December 3, 2019 12:17 To: Sebastien Han <shan@redhat.com>; dev@ceph.io <dev@ceph.io> Cc: Travis Nielsen <tnielsen@redhat.com> Subject: Re: ceph-volume simple disk scenario without LVM for OSD on PVC We could play with making the entrypoing a simple bash script that traps the EXIT signal to kill the captured PID of the OSD. A `wait` might or might not be necessary. Similar thing in practice: http://mywiki.wooledge.org/SignalTrap#When_is_the_signal_handled.3F Blaine ________________________________ From: Sebastien Han <shan@redhat.com> Sent: Tuesday, December 3, 2019 09:55 To: dev@ceph.io <dev@ceph.io> Cc: Travis Nielsen <tnielsen@redhat.com> Subject: ceph-volume simple disk scenario without LVM for OSD on PVC Hi, I've started working on a saner way to deploy OSD with Rook so that they don't use the rook binary image. Why were/are we using the rook binary to activate the OSD? A bit of background on containers first, when executing a container, we need to provide a command entrypoint that will act as PID 1. So if you want to do pre/post action before running the process you need to use a wrapper. In Rook, that's the rook binary, which has a CLI and can then "activate" an OSD. Currently, this "rook osd activate" call does the following: * sed the lvm.conf * run c-v lvm activate * run the osd process On shutdown, we intercept the signal, "kill -9" the osd and de-activate the LV. I have a patch here: https://github.com/rook/rook/pull/4386, that solves the initial bullet points but one thing we cannot do is the signal catching and the lv de-activation. Before you ask, Kubernetes has pre/post-hook but they are not reliable, it's known and documented that there is no guarantee they would actually run before or after the container starts/stops. We tried and we had issues. Why do we want to stop using the rook binary for activation? Because each time we get a new binary version (new operator version), this will restart all the OSDs, even if the deployment spec didn't change, at least if nothing else than the rook image version changed. Also with containers, we have seen so many issues working with LVM, just to name a few: * adapt lvm filters * interactions with udev - need to tune the lvm config, even c-v itself has lvm flag to not sync with udev built-in * several bindmounts * lvm package must be present on the host even if running in containers * SELinux, yes lvm calls SELinux commands under the hood and pollute the logs in some scenarios Currently, one of the ways I can see this working is by not using LVM when bootstrapping OSDs. Unfortunately, some of the logic cannot go in the OSD code since the lv de-activation happens after the OSD stops. We need to de-activate the LV so when running in the Cloud the block can safely be re-attached to a new machine without LVM issues. I know this will be a bit challenging and might ultimately look like ceph-disk but it'd be nice to consider it. What about a small prototype for Bluestore with block/db/wal on the same disk? If this gets rejected, I might try a prototype for not using c-v in Rook or something else that might come up with this discussion. Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Tue, Dec 3, 2019 at 11:56 AM Sebastien Han <shan@redhat.com> wrote:
Hi,
I've started working on a saner way to deploy OSD with Rook so that they don't use the rook binary image.
Why were/are we using the rook binary to activate the OSD?
A bit of background on containers first, when executing a container, we need to provide a command entrypoint that will act as PID 1. So if you want to do pre/post action before running the process you need to use a wrapper. In Rook, that's the rook binary, which has a CLI and can then "activate" an OSD. Currently, this "rook osd activate" call does the following:
* sed the lvm.conf * run c-v lvm activate * run the osd process
On shutdown, we intercept the signal, "kill -9" the osd and de-activate the LV.
I have a patch here: https://github.com/rook/rook/pull/4386, that solves the initial bullet points but one thing we cannot do is the signal catching and the lv de-activation. Before you ask, Kubernetes has pre/post-hook but they are not reliable, it's known and documented that there is no guarantee they would actually run before or after the container starts/stops. We tried and we had issues.
Why do we want to stop using the rook binary for activation? Because each time we get a new binary version (new operator version), this will restart all the OSDs, even if the deployment spec didn't change, at least if nothing else than the rook image version changed.
Also with containers, we have seen so many issues working with LVM, just to name a few:
* adapt lvm filters * interactions with udev - need to tune the lvm config, even c-v itself has lvm flag to not sync with udev built-in * several bindmounts * lvm package must be present on the host even if running in containers * SELinux, yes lvm calls SELinux commands under the hood and pollute the logs in some scenarios
Currently, one of the ways I can see this working is by not using LVM when bootstrapping OSDs. Unfortunately, some of the logic cannot go in the OSD code since the lv de-activation happens after the OSD stops. We need to de-activate the LV so when running in the Cloud the block can safely be re-attached to a new machine without LVM issues.
I know this will be a bit challenging and might ultimately look like ceph-disk but it'd be nice to consider it. What about a small prototype for Bluestore with block/db/wal on the same disk?
You raise some good points here, and I agree that there are many issues with containers and LVM. There were also quite a few issues with ceph-disk in containers, but those issues are not as relevant as making the OSD provisioning easier for everyone else. One of the main ideas I brought up when trying to design ceph-volume was to be completely agnostic on how the OSDs came to be: partitions? full devices? LVM? something else? It was interesting to imagine a scenario where the setup didn't matter much, and ceph-volume would just be in charge of "activating" (ensuring everything is ready for the ceph-osd daemon). That idea got push-back in favor of being opinionated and choosing LVM. The amount of internals ceph-volume has to deal specifically with LVM is enormous, because with LVM came the requests with having more flexibility, and more options to make it easier to use. The `simple` sub-command was an attempt to introduce the hands-off approach to OSD activation, by requiring just a little bit of metadata in /etc/ceph/osd/*.json, where each OSD would represent a single JSON file with some information. That approach not only works well for ceph-disk OSDs, but should also work well with whatever else that you may come up with... have you tried with `simple` and not gotten results? If so, what went wrong? Another option if `simple` doesn't achieve what Rook needs, is perhaps implementing a separate sub-command (ceph-volume container?) that could be implemented as a plugin so that it reuses all the well-tested utilities that ceph-volume already has. The ZFS plugin did something like that already. Creating OSDs on your (Rook's) own is a *very* hard task to get right, not to mention the many different ways OSDs allow you to configure them: filestore (dedicated, collocated), bluestore (data, data+db, data+wal, data+db+wal), dmcrypt or unencrypted. Plus other nuances like talking to the monitor, and sending/retrieving information that has changed between releases.
If this gets rejected, I might try a prototype for not using c-v in Rook or something else that might come up with this discussion.
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood." _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Tue, 3 Dec 2019, Sebastien Han wrote:
Hi,
I've started working on a saner way to deploy OSD with Rook so that they don't use the rook binary image.
Why were/are we using the rook binary to activate the OSD?
A bit of background on containers first, when executing a container, we need to provide a command entrypoint that will act as PID 1. So if you want to do pre/post action before running the process you need to use a wrapper. In Rook, that's the rook binary, which has a CLI and can then "activate" an OSD. Currently, this "rook osd activate" call does the following:
* sed the lvm.conf * run c-v lvm activate * run the osd process
On shutdown, we intercept the signal, "kill -9" the osd and de-activate the LV.
I have a patch here: https://github.com/rook/rook/pull/4386, that solves the initial bullet points but one thing we cannot do is the signal catching and the lv de-activation.
What if we implement a ceph-volume command similar to 'activate' that *also* runs ceph-osd, catches the signal, and cleans up LVM afterwards? It occurs to me that we probably want a similar sequence for ceph-daemon too in the dm-crypt case, where ideally we'd set up the encrypted device, start the osd, and on shutdown, tear it down again.
Before you ask, Kubernetes has pre/post-hook but they are not reliable, it's known and documented that there is no guarantee they would actually run before or after the container starts/stops. We tried and we had issues.
Why do we want to stop using the rook binary for activation? Because each time we get a new binary version (new operator version), this will restart all the OSDs, even if the deployment spec didn't change, at least if nothing else than the rook image version changed.
Also with containers, we have seen so many issues working with LVM, just to name a few:
* adapt lvm filters * interactions with udev - need to tune the lvm config, even c-v itself has lvm flag to not sync with udev built-in * several bindmounts * lvm package must be present on the host even if running in containers * SELinux, yes lvm calls SELinux commands under the hood and pollute the logs in some scenarios
Currently, one of the ways I can see this working is by not using LVM when bootstrapping OSDs. Unfortunately, some of the logic cannot go in the OSD code since the lv de-activation happens after the OSD stops. We need to de-activate the LV so when running in the Cloud the block can safely be re-attached to a new machine without LVM issues.
I know this will be a bit challenging and might ultimately look like ceph-disk but it'd be nice to consider it. What about a small prototype for Bluestore with block/db/wal on the same disk?
If this gets rejected, I might try a prototype for not using c-v in Rook or something else that might come up with this discussion.
An LVM-less approach is appealing. The main case that it doesn't cover is an encrypted device: we need somewhere to stash metadata about the encrypted device and the key that's used to fetch the decryption key. This is very awkward to do with a bare device. I think we'll need something like it eventually for seastore, but I'm worried about building yet another not-quite-as-general-as-we'd-hoped scheme. Something that explicitly does bluestore only and does not support dmcrypt could be pretty straightforward, though... I think it would basically just have to use the bluestore device label to populate a simple .json or /var/lib/ceph/osd directory inside the container. sage
On 2019-12-03T19:58:55, Sage Weil <sage@newdream.net> wrote:
An LVM-less approach is appealing.
I'm not sure. LVM provides many management functions - such as being able to transparently re-map/move data from one device to another, or increasing LV sizes, say - that otherwise would need to be reimplemented by the OSD processes.
Something that explicitly does bluestore only and does not support dmcrypt could be pretty straightforward, though...
Yes, but given the current deployment ratios, I'd bet this would also not be applicable to the majority of deployments. Yes, it'd get the very very simple ones of the ground, but we'd still have to solve the actual problems. (Plus then maintain the "simple" way on top, and how to go from there to the more complex one, feature-disparities, DriveGroups for each, etc etc) -- SUSE Software Solutions Germany GmbH, MD: Felix Imendörffer, HRB 36809 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
I don't think to rehearse the list of features of LVM all the time is a valid argument as we are not even using 10% what LVM is capable of, and everything we do with it could be done with partitions. The only thing I see is a new layer that adds complexity to the setup, a tool that (as Alfredo said, "The amount of internals ceph-volume has to deal specifically with LVM is enormous") spends most its time figuring out how things are layered. I guess I'd be willing to accept LVM a bit more if someone can give me a single feature of LVM that we desperately need and that nothing else has. As of today, I don't think we have demonstrated the value of using LVM instead of block/partitions because, again, all the races we found with ceph-disk were ultimately fixed, and they did not apply to containers! With the adoption and growth of container, all the logic from the host with udev rules and systemd is impossible to replicate without over-engineering our containerized environment. The simple fact that devices are held by LVM (active) is a nightmare to work with when running on PVC in the Cloud, and we are seeing a higher demand for running Rook in the Cloud. Also, the time we spent on tuning lvm flags for containers is ridiculous. Device mobility across host is something we had with ceph-disk, and we lost it with LVM without applying manual commands (to activate/de-activate VG), and now we need it back when running portable OSD in the Cloud... Additionally, we live with too many dependencies on the host (lvm packages needed, lvm systemd), which sounds a bit silly when running on k8s since we (in theory) have no possible interaction with the host configuration. Sorry, this has turned again into a ceph-disk/ceph-volume discussion... Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." On Wed, Dec 4, 2019 at 12:13 PM Lars Marowsky-Bree <lmb@suse.com> wrote:
On 2019-12-03T19:58:55, Sage Weil <sage@newdream.net> wrote:
An LVM-less approach is appealing.
I'm not sure.
LVM provides many management functions - such as being able to transparently re-map/move data from one device to another, or increasing LV sizes, say - that otherwise would need to be reimplemented by the OSD processes.
Something that explicitly does bluestore only and does not support dmcrypt could be pretty straightforward, though...
Yes, but given the current deployment ratios, I'd bet this would also not be applicable to the majority of deployments. Yes, it'd get the very very simple ones of the ground, but we'd still have to solve the actual problems. (Plus then maintain the "simple" way on top, and how to go from there to the more complex one, feature-disparities, DriveGroups for each, etc etc)
-- SUSE Software Solutions Germany GmbH, MD: Felix Imendörffer, HRB 36809 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden) _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Thu, Dec 5, 2019 at 4:34 AM Sebastien Han <shan@redhat.com> wrote:
I don't think to rehearse the list of features of LVM all the time is a valid argument as we are not even using 10% what LVM is capable of, and everything we do with it could be done with partitions. The only thing I see is a new layer that adds complexity to the setup, a tool that (as Alfredo said, "The amount of internals ceph-volume has to deal specifically with LVM is enormous") spends most its time figuring out how things are layered. I guess I'd be willing to accept LVM a bit more if someone can give me a single feature of LVM that we desperately need and that nothing else has.
As of today, I don't think we have demonstrated the value of using LVM instead of block/partitions because, again, all the races we found with ceph-disk were ultimately fixed, and they did not apply to containers!
I would still be interested to see how (why) partitions created by rook and then populated in /etc/ceph/osd/<id>.json does not work for the purposes you've stated. It doesn't tie you to LVM which seems problematic for you, and still allows you to do all the OSD variations including dmcrypt.
With the adoption and growth of container, all the logic from the host with udev rules and systemd is impossible to replicate without over-engineering our containerized environment. The simple fact that devices are held by LVM (active) is a nightmare to work with when running on PVC in the Cloud, and we are seeing a higher demand for running Rook in the Cloud. Also, the time we spent on tuning lvm flags for containers is ridiculous. Device mobility across host is something we had with ceph-disk, and we lost it with LVM without applying manual commands (to activate/de-activate VG), and now we need it back when running portable OSD in the Cloud...
Additionally, we live with too many dependencies on the host (lvm packages needed, lvm systemd), which sounds a bit silly when running on k8s since we (in theory) have no possible interaction with the host configuration. Sorry, this has turned again into a ceph-disk/ceph-volume discussion...
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Wed, Dec 4, 2019 at 12:13 PM Lars Marowsky-Bree <lmb@suse.com> wrote:
On 2019-12-03T19:58:55, Sage Weil <sage@newdream.net> wrote:
An LVM-less approach is appealing.
I'm not sure.
LVM provides many management functions - such as being able to transparently re-map/move data from one device to another, or increasing LV sizes, say - that otherwise would need to be reimplemented by the OSD processes.
Something that explicitly does bluestore only and does not support dmcrypt could be pretty straightforward, though...
Yes, but given the current deployment ratios, I'd bet this would also not be applicable to the majority of deployments. Yes, it'd get the very very simple ones of the ground, but we'd still have to solve the actual problems. (Plus then maintain the "simple" way on top, and how to go from there to the more complex one, feature-disparities, DriveGroups for each, etc etc)
-- SUSE Software Solutions Germany GmbH, MD: Felix Imendörffer, HRB 36809 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden) _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Currently, Rook does not create partitions and this is my last resort option. Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." On Thu, Dec 5, 2019 at 12:35 PM Alfredo Deza <adeza@redhat.com> wrote:
On Thu, Dec 5, 2019 at 4:34 AM Sebastien Han <shan@redhat.com> wrote:
I don't think to rehearse the list of features of LVM all the time is a valid argument as we are not even using 10% what LVM is capable of, and everything we do with it could be done with partitions. The only thing I see is a new layer that adds complexity to the setup, a tool that (as Alfredo said, "The amount of internals ceph-volume has to deal specifically with LVM is enormous") spends most its time figuring out how things are layered. I guess I'd be willing to accept LVM a bit more if someone can give me a single feature of LVM that we desperately need and that nothing else has.
As of today, I don't think we have demonstrated the value of using LVM instead of block/partitions because, again, all the races we found with ceph-disk were ultimately fixed, and they did not apply to containers!
I would still be interested to see how (why) partitions created by rook and then populated in /etc/ceph/osd/<id>.json does not work for the purposes you've stated. It doesn't tie you to LVM which seems problematic for you, and still allows you to do all the OSD variations including dmcrypt.
With the adoption and growth of container, all the logic from the host with udev rules and systemd is impossible to replicate without over-engineering our containerized environment. The simple fact that devices are held by LVM (active) is a nightmare to work with when running on PVC in the Cloud, and we are seeing a higher demand for running Rook in the Cloud. Also, the time we spent on tuning lvm flags for containers is ridiculous. Device mobility across host is something we had with ceph-disk, and we lost it with LVM without applying manual commands (to activate/de-activate VG), and now we need it back when running portable OSD in the Cloud...
Additionally, we live with too many dependencies on the host (lvm packages needed, lvm systemd), which sounds a bit silly when running on k8s since we (in theory) have no possible interaction with the host configuration. Sorry, this has turned again into a ceph-disk/ceph-volume discussion...
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Wed, Dec 4, 2019 at 12:13 PM Lars Marowsky-Bree <lmb@suse.com> wrote:
On 2019-12-03T19:58:55, Sage Weil <sage@newdream.net> wrote:
An LVM-less approach is appealing.
I'm not sure.
LVM provides many management functions - such as being able to transparently re-map/move data from one device to another, or increasing LV sizes, say - that otherwise would need to be reimplemented by the OSD processes.
Something that explicitly does bluestore only and does not support dmcrypt could be pretty straightforward, though...
Yes, but given the current deployment ratios, I'd bet this would also not be applicable to the majority of deployments. Yes, it'd get the very very simple ones of the ground, but we'd still have to solve the actual problems. (Plus then maintain the "simple" way on top, and how to go from there to the more complex one, feature-disparities, DriveGroups for each, etc etc)
-- SUSE Software Solutions Germany GmbH, MD: Felix Imendörffer, HRB 36809 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden) _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On 05.12.19 10:33, Sebastien Han wrote:
Sorry, this has turned again into a ceph-disk/ceph-volume discussion...
Yes because that's basically what you're doing here. Personally I don't see the value in just switching back again as I didn't found real good and absolutely necessary reasons to do so and why rook and/or ceph-volume couldn't be fixed in that regard. I also didn't find out why the new solution should be any better in the case that we don't talk about switching back in a few month again because all of those issues are fixed now - You said it yourself "ultimately fixed" which at the same time means that such tools need some time to reach the point were all those issues are fixed. We need to start to talk across boundaries and to reach out and involve other project maintainer to work on real solutions instead of moving out of their way. Just imagine we would switch back to partitions, that would mean that we're doing again a breaking change between releases. This change would result in the same thing such like filestore->bluestore or ceph-disk->ceph-volume - people have to rewrite the whole data of their cluster at one point or another (or do we want to support both tools forever?). I have the feeling that we sometimes treat this whole project still just like a kick-starter project were everything can be switched and changed between releases in any direction and whenever we think it would be nice to do so. Could we please start to think about that there's a big and luckily growing user base behind that are actively relaying on this solution and they're not keen on moving their data around just because we had a "feeling"? No one should feel personally offended by the last remark but introducing such changes isn't something we could/should just do. We really should have 100% valid arguments why replacing one tool by another tool, which we had already in the past, now magically solves everything and we're sure that we can't fix such issues in the current solution. Kai -- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
Inline: Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." On Fri, Dec 6, 2019 at 12:22 AM Kai Wagner <kwagner@suse.com> wrote:
On 05.12.19 10:33, Sebastien Han wrote:
Sorry, this has turned again into a ceph-disk/ceph-volume discussion...
Yes because that's basically what you're doing here.
So what?
Personally I don't see the value in just switching back again as I didn't found real good and absolutely necessary reasons to do so and why rook and/or ceph-volume couldn't be fixed in that regard. I also didn't find out why the new solution should be any better in the case that we don't talk about switching back in a few month again because all of those issues are fixed now - You said it yourself "ultimately fixed" which at the same time means that such tools need some time to reach the point were all those issues are fixed. We need to start to talk across boundaries and to reach out and involve other project maintainer to work on real solutions instead of moving out of their way.
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
Just imagine we would switch back to partitions, that would mean that we're doing again a breaking change between releases. This change would result in the same thing such like filestore->bluestore or ceph-disk->ceph-volume - people have to rewrite the whole data of their cluster at one point or another (or do we want to support both tools forever?).
That's completely wrong. It's not because you change the tool to provision OSDs that you must trash and re-create your OSD. That's why ceph-volume has the "simple scan" command interface. ceph-ansible has handled that migration pretty well transparently without any breaking change.
I have the feeling that we sometimes treat this whole project still just like a kick-starter project were everything can be switched and changed between releases in any direction and whenever we think it would be nice to do so. Could we please start to think about that there's a big and luckily growing user base behind that are actively relaying on this solution and they're not keen on moving their data around just because we had a "feeling"?
I disagree, we are just talking about the provisioning tool, no breaking changes involved and I don't even recall any major re-architecture that ever introduced difficulties to users.
No one should feel personally offended by the last remark but introducing such changes isn't something we could/should just do. We really should have 100% valid arguments why replacing one tool by another tool, which we had already in the past, now magically solves everything and we're sure that we can't fix such issues in the current solution.
100% is a myth, there is no such thing and no we didn't have that in the past. Back then, I warned about what was coming with containers and here we are... Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with Anyway, again, that discussion has gotten intense and we should probably stop the storm here. I'm not a big fan of the "ceph-volume activate" wrapper idea but short-term that's probably the best thing we can do, I'll send an e-mail for this shortly. I'll also investigate not using ceph-volume in Rook when it comes to collocating block/db/wal in the same disk for a bluestore OSD. Thanks everyone for your replies, appreciate the inputs.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
Inline:
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Fri, Dec 6, 2019 at 12:22 AM Kai Wagner <kwagner@suse.com> wrote:
On 05.12.19 10:33, Sebastien Han wrote:
Sorry, this has turned again into a ceph-disk/ceph-volume discussion...
Yes because that's basically what you're doing here.
So what?
Personally I don't see the value in just switching back again as I didn't found real good and absolutely necessary reasons to do so and why rook and/or ceph-volume couldn't be fixed in that regard. I also didn't find out why the new solution should be any better in the case that we don't talk about switching back in a few month again because all of those issues are fixed now - You said it yourself "ultimately fixed" which at the same time means that such tools need some time to reach the point were all those issues are fixed. We need to start to talk across boundaries and to reach out and involve other project maintainer to work on real solutions instead of moving out of their way.
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again. I don't think this was the argument. The intention is to at least try and fix
On Fri, Dec 06, 2019 at 10:00:04AM +0100, Sebastien Han wrote: the code we have, instead of inventing something new yet again. I'm still not sure what you mean with "a lot of pain". Nothing ever made it into the c-v tracker to see if things can be improved. What are these painful issue, are they so bad that we can not work on those as a community? I'm aware of https://github.com/rook/rook/pull/3771 and https://github.com/rook/rook/pull/4219. And again I'm more than happy to improve the c-v side of things, I just can't when I'm not aware.
Just imagine we would switch back to partitions, that would mean that we're doing again a breaking change between releases. This change would result in the same thing such like filestore->bluestore or ceph-disk->ceph-volume - people have to rewrite the whole data of their cluster at one point or another (or do we want to support both tools forever?).
That's completely wrong. It's not because you change the tool to provision OSDs that you must trash and re-create your OSD. That's why ceph-volume has the "simple scan" command interface. ceph-ansible has handled that migration pretty well transparently without any breaking change.
So you propose user with lvm deployed OSDs can migrate without redeploying or any service interruption? How would that work?
I have the feeling that we sometimes treat this whole project still just like a kick-starter project were everything can be switched and changed between releases in any direction and whenever we think it would be nice to do so. Could we please start to think about that there's a big and luckily growing user base behind that are actively relaying on this solution and they're not keen on moving their data around just because we had a "feeling"?
I disagree, we are just talking about the provisioning tool, no breaking changes involved and I don't even recall any major re-architecture that ever introduced difficulties to users.
Every update usually introduces pain points for users. Bluestore was certainly one. Sure no user has to migrate, only if they want the shiny new features.
No one should feel personally offended by the last remark but introducing such changes isn't something we could/should just do. We really should have 100% valid arguments why replacing one tool by another tool, which we had already in the past, now magically solves everything and we're sure that we can't fix such issues in the current solution.
100% is a myth, there is no such thing and no we didn't have that in the past. Back then, I warned about what was coming with containers and here we are... Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Anyway, again, that discussion has gotten intense and we should probably stop the storm here. I'm not a big fan of the "ceph-volume activate" wrapper idea but short-term that's probably the best thing we can do, I'll send an e-mail for this shortly. I'll also investigate not using ceph-volume in Rook when it comes to collocating block/db/wal in the same disk for a bluestore OSD.
Thanks everyone for your replies, appreciate the inputs.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
-- Jan Fajerski Senior Software Engineer Enterprise Storage SUSE Software Solutions Germany GmbH Maxfeldstr. 5, 90409 Nürnberg, Germany (HRB 36809, AG Nürnberg) Geschäftsführer: Felix Imendörffer
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." On Fri, Dec 6, 2019 at 10:41 AM Jan Fajerski <jfajerski@suse.com> wrote:
Inline:
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Fri, Dec 6, 2019 at 12:22 AM Kai Wagner <kwagner@suse.com> wrote:
On 05.12.19 10:33, Sebastien Han wrote:
Sorry, this has turned again into a ceph-disk/ceph-volume discussion...
Yes because that's basically what you're doing here.
So what?
Personally I don't see the value in just switching back again as I didn't found real good and absolutely necessary reasons to do so and why rook and/or ceph-volume couldn't be fixed in that regard. I also didn't find out why the new solution should be any better in the case that we don't talk about switching back in a few month again because all of those issues are fixed now - You said it yourself "ultimately fixed" which at the same time means that such tools need some time to reach the point were all those issues are fixed. We need to start to talk across boundaries and to reach out and involve other project maintainer to work on real solutions instead of moving out of their way.
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again. I don't think this was the argument. The intention is to at least try and fix
On Fri, Dec 06, 2019 at 10:00:04AM +0100, Sebastien Han wrote: the code we have, instead of inventing something new yet again. I'm still not sure what you mean with "a lot of pain". Nothing ever made it into the c-v tracker to see if things can be improved. What are these painful issue, are they so bad that we can not work on those as a community? I'm aware of https://github.com/rook/rook/pull/3771 and https://github.com/rook/rook/pull/4219. And again I'm more than happy to improve the c-v side of things, I just can't when I'm not aware.
There are certainly more than the 2 you mentioned, I can build up a list of you insist. Some of them were worked-around in Rook directly (due to time constraint), some are pure container permission related because of LVM (e,g. bindmounts). Yes I know you're always willing to help and I appreciate that. As mentioned in the previous reply I'll send a new email to propose a solution that should be implemented in c-v then we will follow up with a tracker bug.
Just imagine we would switch back to partitions, that would mean that we're doing again a breaking change between releases. This change would result in the same thing such like filestore->bluestore or ceph-disk->ceph-volume - people have to rewrite the whole data of their cluster at one point or another (or do we want to support both tools forever?).
That's completely wrong. It's not because you change the tool to provision OSDs that you must trash and re-create your OSD. That's why ceph-volume has the "simple scan" command interface. ceph-ansible has handled that migration pretty well transparently without any breaking change.
So you propose user with lvm deployed OSDs can migrate without redeploying or any service interruption? How would that work?
There is nothing to migrate, people with ceph-disk-based OSDs never migrated. Again, see "simple scan" + "activate" combo, that's how things continue to operate.
I have the feeling that we sometimes treat this whole project still just like a kick-starter project were everything can be switched and changed between releases in any direction and whenever we think it would be nice to do so. Could we please start to think about that there's a big and luckily growing user base behind that are actively relaying on this solution and they're not keen on moving their data around just because we had a "feeling"?
I disagree, we are just talking about the provisioning tool, no breaking changes involved and I don't even recall any major re-architecture that ever introduced difficulties to users.
Every update usually introduces pain points for users. Bluestore was certainly one. Sure no user has to migrate, only if they want the shiny new features.
No one should feel personally offended by the last remark but introducing such changes isn't something we could/should just do. We really should have 100% valid arguments why replacing one tool by another tool, which we had already in the past, now magically solves everything and we're sure that we can't fix such issues in the current solution.
100% is a myth, there is no such thing and no we didn't have that in the past. Back then, I warned about what was coming with containers and here we are... Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Anyway, again, that discussion has gotten intense and we should probably stop the storm here. I'm not a big fan of the "ceph-volume activate" wrapper idea but short-term that's probably the best thing we can do, I'll send an e-mail for this shortly. I'll also investigate not using ceph-volume in Rook when it comes to collocating block/db/wal in the same disk for a bluestore OSD.
Thanks everyone for your replies, appreciate the inputs.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
-- Jan Fajerski Senior Software Engineer Enterprise Storage SUSE Software Solutions Germany GmbH Maxfeldstr. 5, 90409 Nürnberg, Germany (HRB 36809, AG Nürnberg) Geschäftsführer: Felix Imendörffer _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Hi Sebastien and thanks for your feedback. On 06.12.19 10:00, Sebastien Han wrote:
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
The reason I disagree here is the scenario were the WAL/DB is on a separate device and a single OSD crashes. In that case you would like to recreate just that single OSD instead of the whole group. Also if we deprecate a tool such like we did with ceph-disk, users have to migrate sooner or later if they don't want to do everything manually on the CLI (by that I mean via fdisk/pure lvm commands and so on). We could argue now that this can still be done on the command line manually but all our efforts are towards simplicity/automation and having everything in the Dashboard. If the underlying tool/functionality isn't there anymore, that isn't possible.
Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Which then leads me to, why couldn't such functionality be implemented into a single tool instead of having two at the end? So don't get me wrong, I'm not saying that I'm against everything I'm just saying that I think this is a topic that should be discussed in more depth. As said, just my two cents here. Kai -- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
Hi Kai, Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." On Fri, Dec 6, 2019 at 10:44 AM Kai Wagner <kwagner@suse.com> wrote:
Hi Sebastien and thanks for your feedback.
On 06.12.19 10:00, Sebastien Han wrote:
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
The reason I disagree here is the scenario were the WAL/DB is on a separate device and a single OSD crashes. In that case you would like to recreate just that single OSD instead of the whole group. Also if we deprecate a tool such like we did with ceph-disk, users have to migrate sooner or later if they don't want to do everything manually on the CLI (by that I mean via fdisk/pure lvm commands and so on).
We could argue now that this can still be done on the command line manually but all our efforts are towards simplicity/automation and having everything in the Dashboard. If the underlying tool/functionality isn't there anymore, that isn't possible.
I understand your position, yes when we start separating block/db/wal things get really complex that's why I'm sticking with block/db/wal in the same block. Also, we haven't seen any request for separating those when running OSDs on PVC in the Cloud. So we would likely continue to do so for a while.
Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Which then leads me to, why couldn't such functionality be implemented into a single tool instead of having two at the end?
So don't get me wrong, I'm not saying that I'm against everything I'm just saying that I think this is a topic that should be discussed in more depth.
Yes, that's for sure.
As said, just my two cents here.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
My thoughts here are still pretty inconclusive... I agree that we should invest a non-LVM mode, but there isn't a way to do that currently that supports dm-crypt that isn't complicated and convoluted, so it cannot be a full replacement for the LVM mode. At the same time, Real Soon Now we're going to be building crimson OSDs backed by ZNS SSDs (and eventually persistent memory), which will also very clearly not be LVM-based. I'm a bit hesitant to introduce a bare-bones bluestore mode right now just because we'll be adding yet another variation soon, and it may be that we construct a general approach to both... but probably not. And the whole point of c-v's architecture was to be pluggable. So maybe a bare-bones bluestore mode makes sense. In the simple case, it really should be *very* simple. But its scope pretty quickly expodes: what about wal and db devices? We have labels for those, so we could support those, also easily... if the user has to partition the devices beforehand manually. They'll immediately want to use the new auto/batch thing, but that's tied to the LVM implementation. And what if one of the db/wal/main devices is an LV and another is not? We'd need to make sure the lvm mode machinery doesn't trigger unless all of its labels are there, but it might be confusing. All of which means that this is probably *only* useful for single-device OSDs. On the one hand, those are increasingly common (hello, all-SSD clusters), but on the other hand, for fast SSDs we may want to deploy N of them per device. Since we can't cover all of that, and at a minimum, we can't cover dm-crypt, Rook will need to behave with the lvm mode one way or another. So we need to have a wrapper (or something similar) no matter what. So I suggest we start there. sage On Fri, 6 Dec 2019, Sebastien Han wrote:
Hi Kai,
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Fri, Dec 6, 2019 at 10:44 AM Kai Wagner <kwagner@suse.com> wrote:
Hi Sebastien and thanks for your feedback.
On 06.12.19 10:00, Sebastien Han wrote:
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
The reason I disagree here is the scenario were the WAL/DB is on a separate device and a single OSD crashes. In that case you would like to recreate just that single OSD instead of the whole group. Also if we deprecate a tool such like we did with ceph-disk, users have to migrate sooner or later if they don't want to do everything manually on the CLI (by that I mean via fdisk/pure lvm commands and so on).
We could argue now that this can still be done on the command line manually but all our efforts are towards simplicity/automation and having everything in the Dashboard. If the underlying tool/functionality isn't there anymore, that isn't possible.
I understand your position, yes when we start separating block/db/wal things get really complex that's why I'm sticking with block/db/wal in the same block. Also, we haven't seen any request for separating those when running OSDs on PVC in the Cloud. So we would likely continue to do so for a while.
Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Which then leads me to, why couldn't such functionality be implemented into a single tool instead of having two at the end?
So don't get me wrong, I'm not saying that I'm against everything I'm just saying that I think this is a topic that should be discussed in more depth.
Yes, that's for sure.
As said, just my two cents here.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
With Bluestore I 've seen more users collocating everything on the same device as the performance is good enough already. Also, this is something that we already demonstrated in various benchmarks. So the block/db/wal on the same device probably suits most of the users out there and drastically reduces the complexity of the setup, as well as increasing availability (if you lose the db/wal device you lose all the osds associated to it and a lot of people are not ready to take that risk). That's why I think this all-in-one mode osd should be simple and robust without another extra LVM layer on it. Even if we change the osd store, this will likely continue to be the same. For more advanced setup, we can keep LVM for now I suppose... Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect "Always give 100%. Unless you're giving blood." On Fri, Dec 6, 2019 at 2:31 PM Sage Weil <sage@newdream.net> wrote:
My thoughts here are still pretty inconclusive...
I agree that we should invest a non-LVM mode, but there isn't a way to do that currently that supports dm-crypt that isn't complicated and convoluted, so it cannot be a full replacement for the LVM mode.
At the same time, Real Soon Now we're going to be building crimson OSDs backed by ZNS SSDs (and eventually persistent memory), which will also very clearly not be LVM-based. I'm a bit hesitant to introduce a bare-bones bluestore mode right now just because we'll be adding yet another variation soon, and it may be that we construct a general approach to both... but probably not. And the whole point of c-v's architecture was to be pluggable.
So maybe a bare-bones bluestore mode makes sense. In the simple case, it really should be *very* simple. But its scope pretty quickly expodes: what about wal and db devices? We have labels for those, so we could support those, also easily... if the user has to partition the devices beforehand manually. They'll immediately want to use the new auto/batch thing, but that's tied to the LVM implementation. And what if one of the db/wal/main devices is an LV and another is not? We'd need to make sure the lvm mode machinery doesn't trigger unless all of its labels are there, but it might be confusing. All of which means that this is probably *only* useful for single-device OSDs. On the one hand, those are increasingly common (hello, all-SSD clusters), but on the other hand, for fast SSDs we may want to deploy N of them per device.
Since we can't cover all of that, and at a minimum, we can't cover dm-crypt, Rook will need to behave with the lvm mode one way or another. So we need to have a wrapper (or something similar) no matter what. So I suggest we start there.
sage
On Fri, 6 Dec 2019, Sebastien Han wrote:
Hi Kai,
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Fri, Dec 6, 2019 at 10:44 AM Kai Wagner <kwagner@suse.com> wrote:
Hi Sebastien and thanks for your feedback.
On 06.12.19 10:00, Sebastien Han wrote:
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
The reason I disagree here is the scenario were the WAL/DB is on a separate device and a single OSD crashes. In that case you would like to recreate just that single OSD instead of the whole group. Also if we deprecate a tool such like we did with ceph-disk, users have to migrate sooner or later if they don't want to do everything manually on the CLI (by that I mean via fdisk/pure lvm commands and so on).
We could argue now that this can still be done on the command line manually but all our efforts are towards simplicity/automation and having everything in the Dashboard. If the underlying tool/functionality isn't there anymore, that isn't possible.
I understand your position, yes when we start separating block/db/wal things get really complex that's why I'm sticking with block/db/wal in the same block. Also, we haven't seen any request for separating those when running OSDs on PVC in the Cloud. So we would likely continue to do so for a while.
Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Which then leads me to, why couldn't such functionality be implemented into a single tool instead of having two at the end?
So don't get me wrong, I'm not saying that I'm against everything I'm just saying that I think this is a topic that should be discussed in more depth.
Yes, that's for sure.
As said, just my two cents here.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Fri, Dec 6, 2019 at 8:31 AM Sage Weil <sage@newdream.net> wrote:
My thoughts here are still pretty inconclusive...
I agree that we should invest a non-LVM mode, but there isn't a way to do that currently that supports dm-crypt that isn't complicated and convoluted, so it cannot be a full replacement for the LVM mode.
The `ceph-volume simple` sub-command does allow dmcrypt. The key is stored in the JSON file in /etc/ceph/osd. Is there a scenario you've seen where this is not possible? The `simple` sub-command would even allow partitions (regardless of ceph-disk).
At the same time, Real Soon Now we're going to be building crimson OSDs backed by ZNS SSDs (and eventually persistent memory), which will also very clearly not be LVM-based. I'm a bit hesitant to introduce a bare-bones bluestore mode right now just because we'll be adding yet another variation soon, and it may be that we construct a general approach to both... but probably not. And the whole point of c-v's architecture was to be pluggable.
So maybe a bare-bones bluestore mode makes sense. In the simple case, it really should be *very* simple. But its scope pretty quickly expodes: what about wal and db devices? We have labels for those, so we could support those, also easily... if the user has to partition the devices beforehand manually. They'll immediately want to use the new auto/batch thing, but that's tied to the LVM implementation. And what if one of the db/wal/main devices is an LV and another is not? We'd need to make sure the lvm mode machinery doesn't trigger unless all of its labels are there, but it might be confusing. All of which means that this is probably *only* useful for single-device OSDs. On the one hand, those are increasingly common (hello, all-SSD clusters), but on the other hand, for fast SSDs we may want to deploy N of them per device.
Since we can't cover all of that, and at a minimum, we can't cover dm-crypt, Rook will need to behave with the lvm mode one way or another. So we need to have a wrapper (or something similar) no matter what. So I suggest we start there.
sage
On Fri, 6 Dec 2019, Sebastien Han wrote:
Hi Kai,
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Fri, Dec 6, 2019 at 10:44 AM Kai Wagner <kwagner@suse.com> wrote:
Hi Sebastien and thanks for your feedback.
On 06.12.19 10:00, Sebastien Han wrote:
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
The reason I disagree here is the scenario were the WAL/DB is on a separate device and a single OSD crashes. In that case you would like to recreate just that single OSD instead of the whole group. Also if we deprecate a tool such like we did with ceph-disk, users have to migrate sooner or later if they don't want to do everything manually on the CLI (by that I mean via fdisk/pure lvm commands and so on).
We could argue now that this can still be done on the command line manually but all our efforts are towards simplicity/automation and having everything in the Dashboard. If the underlying tool/functionality isn't there anymore, that isn't possible.
I understand your position, yes when we start separating block/db/wal things get really complex that's why I'm sticking with block/db/wal in the same block. Also, we haven't seen any request for separating those when running OSDs on PVC in the Cloud. So we would likely continue to do so for a while.
Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Which then leads me to, why couldn't such functionality be implemented into a single tool instead of having two at the end?
So don't get me wrong, I'm not saying that I'm against everything I'm just saying that I think this is a topic that should be discussed in more depth.
Yes, that's for sure.
As said, just my two cents here.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Fri, 6 Dec 2019, Alfredo Deza wrote:
On Fri, Dec 6, 2019 at 8:31 AM Sage Weil <sage@newdream.net> wrote:
My thoughts here are still pretty inconclusive...
I agree that we should invest a non-LVM mode, but there isn't a way to do that currently that supports dm-crypt that isn't complicated and convoluted, so it cannot be a full replacement for the LVM mode.
The `ceph-volume simple` sub-command does allow dmcrypt. The key is stored in the JSON file in /etc/ceph/osd.
Is there a scenario you've seen where this is not possible? The `simple` sub-command would even allow partitions (regardless of ceph-disk).
For the dm-crypt case, I'm assuming we need the key to be attached to the device in some way. LVM does this with another LV (IIRC); ceph-disk did it with a tiny partition. Putting it in /etc/ceph means you can't move a disk to another server without manually copying files arounds. sage
At the same time, Real Soon Now we're going to be building crimson OSDs backed by ZNS SSDs (and eventually persistent memory), which will also very clearly not be LVM-based. I'm a bit hesitant to introduce a bare-bones bluestore mode right now just because we'll be adding yet another variation soon, and it may be that we construct a general approach to both... but probably not. And the whole point of c-v's architecture was to be pluggable.
So maybe a bare-bones bluestore mode makes sense. In the simple case, it really should be *very* simple. But its scope pretty quickly expodes: what about wal and db devices? We have labels for those, so we could support those, also easily... if the user has to partition the devices beforehand manually. They'll immediately want to use the new auto/batch thing, but that's tied to the LVM implementation. And what if one of the db/wal/main devices is an LV and another is not? We'd need to make sure the lvm mode machinery doesn't trigger unless all of its labels are there, but it might be confusing. All of which means that this is probably *only* useful for single-device OSDs. On the one hand, those are increasingly common (hello, all-SSD clusters), but on the other hand, for fast SSDs we may want to deploy N of them per device.
Since we can't cover all of that, and at a minimum, we can't cover dm-crypt, Rook will need to behave with the lvm mode one way or another. So we need to have a wrapper (or something similar) no matter what. So I suggest we start there.
sage
On Fri, 6 Dec 2019, Sebastien Han wrote:
Hi Kai,
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Fri, Dec 6, 2019 at 10:44 AM Kai Wagner <kwagner@suse.com> wrote:
Hi Sebastien and thanks for your feedback.
On 06.12.19 10:00, Sebastien Han wrote:
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
The reason I disagree here is the scenario were the WAL/DB is on a separate device and a single OSD crashes. In that case you would like to recreate just that single OSD instead of the whole group. Also if we deprecate a tool such like we did with ceph-disk, users have to migrate sooner or later if they don't want to do everything manually on the CLI (by that I mean via fdisk/pure lvm commands and so on).
We could argue now that this can still be done on the command line manually but all our efforts are towards simplicity/automation and having everything in the Dashboard. If the underlying tool/functionality isn't there anymore, that isn't possible.
I understand your position, yes when we start separating block/db/wal things get really complex that's why I'm sticking with block/db/wal in the same block. Also, we haven't seen any request for separating those when running OSDs on PVC in the Cloud. So we would likely continue to do so for a while.
Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Which then leads me to, why couldn't such functionality be implemented into a single tool instead of having two at the end?
So don't get me wrong, I'm not saying that I'm against everything I'm just saying that I think this is a topic that should be discussed in more depth.
Yes, that's for sure.
As said, just my two cents here.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
The key in /etc/ceph also gets lost if you repave the OS. I discovered the hard way when I had to repave a node and the OSDs wouldn’t start. They had to be redeployed, which was a whole lot of data movement that shouldn’t have needed to happen. I subsequently found another 150 or so that I redeployed to avoid another rude awakening. The current (well as of Luminous) scheme is to store them on the mons, is that somehow no longer feasible?
On Fri, 6 Dec 2019, Alfredo Deza wrote:
On Fri, Dec 6, 2019 at 8:31 AM Sage Weil <sage@newdream.net> wrote:
My thoughts here are still pretty inconclusive...
I agree that we should invest a non-LVM mode, but there isn't a way to do that currently that supports dm-crypt that isn't complicated and convoluted, so it cannot be a full replacement for the LVM mode.
The `ceph-volume simple` sub-command does allow dmcrypt. The key is stored in the JSON file in /etc/ceph/osd.
Is there a scenario you've seen where this is not possible? The `simple` sub-command would even allow partitions (regardless of ceph-disk).
For the dm-crypt case, I'm assuming we need the key to be attached to the device in some way. LVM does this with another LV (IIRC); ceph-disk did it with a tiny partition. Putting it in /etc/ceph means you can't move a disk to another server without manually copying files arounds.
sage
On Fri, 6 Dec 2019, Anthony D'Atri wrote:
The key in /etc/ceph also gets lost if you repave the OS. I discovered the hard way when I had to repave a node and the OSDs wouldn’t start. They had to be redeployed, which was a whole lot of data movement that shouldn’t have needed to happen. I subsequently found another 150 or so that I redeployed to avoid another rude awakening.
The current (well as of Luminous) scheme is to store them on the mons, is that somehow no longer feasible?
Yeah, that's the scheme I'm talking about. There is metadata and a key-fetching-key on the device itself, though, that is used to get the actual dm-crypt decryption key from the mon. That needs to go somewhere (another partition, a small LV, LVM metadata, etc.). The c-v simple mode is more like the pre-luminous scheme, so not really useful for dm-crypt. sage
On Fri, 6 Dec 2019, Alfredo Deza wrote:
On Fri, Dec 6, 2019 at 8:31 AM Sage Weil <sage@newdream.net> wrote:
My thoughts here are still pretty inconclusive...
I agree that we should invest a non-LVM mode, but there isn't a way to do that currently that supports dm-crypt that isn't complicated and convoluted, so it cannot be a full replacement for the LVM mode.
The `ceph-volume simple` sub-command does allow dmcrypt. The key is stored in the JSON file in /etc/ceph/osd.
Is there a scenario you've seen where this is not possible? The `simple` sub-command would even allow partitions (regardless of ceph-disk).
For the dm-crypt case, I'm assuming we need the key to be attached to the device in some way. LVM does this with another LV (IIRC); ceph-disk did it with a tiny partition. Putting it in /etc/ceph means you can't move a disk to another server without manually copying files arounds.
sage
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Fri, Dec 6, 2019 at 9:39 AM Sage Weil <sage@newdream.net> wrote:
On Fri, 6 Dec 2019, Alfredo Deza wrote:
On Fri, Dec 6, 2019 at 8:31 AM Sage Weil <sage@newdream.net> wrote:
My thoughts here are still pretty inconclusive...
I agree that we should invest a non-LVM mode, but there isn't a way to do that currently that supports dm-crypt that isn't complicated and convoluted, so it cannot be a full replacement for the LVM mode.
The `ceph-volume simple` sub-command does allow dmcrypt. The key is stored in the JSON file in /etc/ceph/osd.
Is there a scenario you've seen where this is not possible? The `simple` sub-command would even allow partitions (regardless of ceph-disk).
For the dm-crypt case, I'm assuming we need the key to be attached to the device in some way.
We don't need it attached to the device. It just happens that ceph-disk has a partition where it would store it and ceph-volume is able to "scan" it and retrieve it so that it can save it in the JSON file.
LVM does this with another LV (IIRC); ceph-disk did it with a tiny partition. Putting it in /etc/ceph means you can't move a disk to another server without manually copying files arounds.
You are right about LVM and ceph-disk. What I'm trying to make clear is: it is entirely possible, and supported, to have an encrypted OSD via the `simple` sub-command in the JSON file. If you need to move disks around this will not work (would love not to even support this at all, as it is an optimization for small clusters) If the OS dies with the keys, then yes, you would need to have those files replaced somehow. In the case of containers, the files are already somewhere else (the host), and in the specific case of Rook, I want to reiterate: it is possible, and supported to have encrypted OSDs with the key in the JSON file.
sage
At the same time, Real Soon Now we're going to be building crimson OSDs backed by ZNS SSDs (and eventually persistent memory), which will also very clearly not be LVM-based. I'm a bit hesitant to introduce a bare-bones bluestore mode right now just because we'll be adding yet another variation soon, and it may be that we construct a general approach to both... but probably not. And the whole point of c-v's architecture was to be pluggable.
So maybe a bare-bones bluestore mode makes sense. In the simple case, it really should be *very* simple. But its scope pretty quickly expodes: what about wal and db devices? We have labels for those, so we could support those, also easily... if the user has to partition the devices beforehand manually. They'll immediately want to use the new auto/batch thing, but that's tied to the LVM implementation. And what if one of the db/wal/main devices is an LV and another is not? We'd need to make sure the lvm mode machinery doesn't trigger unless all of its labels are there, but it might be confusing. All of which means that this is probably *only* useful for single-device OSDs. On the one hand, those are increasingly common (hello, all-SSD clusters), but on the other hand, for fast SSDs we may want to deploy N of them per device.
Since we can't cover all of that, and at a minimum, we can't cover dm-crypt, Rook will need to behave with the lvm mode one way or another. So we need to have a wrapper (or something similar) no matter what. So I suggest we start there.
sage
On Fri, 6 Dec 2019, Sebastien Han wrote:
Hi Kai,
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Fri, Dec 6, 2019 at 10:44 AM Kai Wagner <kwagner@suse.com> wrote:
Hi Sebastien and thanks for your feedback.
On 06.12.19 10:00, Sebastien Han wrote:
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
The reason I disagree here is the scenario were the WAL/DB is on a separate device and a single OSD crashes. In that case you would like to recreate just that single OSD instead of the whole group. Also if we deprecate a tool such like we did with ceph-disk, users have to migrate sooner or later if they don't want to do everything manually on the CLI (by that I mean via fdisk/pure lvm commands and so on).
We could argue now that this can still be done on the command line manually but all our efforts are towards simplicity/automation and having everything in the Dashboard. If the underlying tool/functionality isn't there anymore, that isn't possible.
I understand your position, yes when we start separating block/db/wal things get really complex that's why I'm sticking with block/db/wal in the same block. Also, we haven't seen any request for separating those when running OSDs on PVC in the Cloud. So we would likely continue to do so for a while.
Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Which then leads me to, why couldn't such functionality be implemented into a single tool instead of having two at the end?
So don't get me wrong, I'm not saying that I'm against everything I'm just saying that I think this is a topic that should be discussed in more depth.
Yes, that's for sure.
As said, just my two cents here.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Fri, Dec 06, 2019 at 01:31:05PM +0000, Sage Weil wrote:
My thoughts here are still pretty inconclusive...
I agree that we should invest a non-LVM mode, but there isn't a way to do that currently that supports dm-crypt that isn't complicated and convoluted, so it cannot be a full replacement for the LVM mode.
At the same time, Real Soon Now we're going to be building crimson OSDs backed by ZNS SSDs (and eventually persistent memory), which will also very clearly not be LVM-based. I'm a bit hesitant to introduce a bare-bones bluestore mode right now just because we'll be adding yet another variation soon, and it may be that we construct a general approach to both... but probably not. And the whole point of c-v's architecture was to be pluggable.
So maybe a bare-bones bluestore mode makes sense. In the simple case, it really should be *very* simple. But its scope pretty quickly expodes: what about wal and db devices? We have labels for those, so we could support those, also easily... if the user has to partition the devices beforehand manually. They'll immediately want to use the new auto/batch thing, but that's tied to the LVM implementation. And what if one of the db/wal/main devices is an LV and another is not? We'd need to make sure the lvm mode machinery doesn't trigger unless all of its labels are there, but it might be confusing. All of which means that this is probably *only* useful for single-device OSDs. On the one hand, those are increasingly common (hello, all-SSD clusters), but on the other hand, for fast SSDs we may want to deploy N of them per device. I don't think keeping a simple or barebones approach will survive contact with real-world deployments. Imho if we want a raw mode, we better be prepared to deal with multi-device OSDs and multi-OSD devices and the partitioning this requires.
Since we can't cover all of that, and at a minimum, we can't cover dm-crypt, Rook will need to behave with the lvm mode one way or another. So we need to have a wrapper (or something similar) no matter what. So I suggest we start there. Agreed.
sage
On Fri, 6 Dec 2019, Sebastien Han wrote:
Hi Kai,
Thanks! ––––––––– Sébastien Han Senior Principal Software Engineer, Storage Architect
"Always give 100%. Unless you're giving blood."
On Fri, Dec 6, 2019 at 10:44 AM Kai Wagner <kwagner@suse.com> wrote:
Hi Sebastien and thanks for your feedback.
On 06.12.19 10:00, Sebastien Han wrote:
ceph-volume is a sunk cost! And your argument basically falls into that paradigm, "oh we have invested so much already, that we cannot stop and we should continue even though this will only bring more trouble". Incapable of accepting this sunk cost. All the issues that have been fixed with a lot of pain. All that pain could have been avoided if LVM wasn't there and pursuing in that direction will only lead us to more pain again.
The reason I disagree here is the scenario were the WAL/DB is on a separate device and a single OSD crashes. In that case you would like to recreate just that single OSD instead of the whole group. Also if we deprecate a tool such like we did with ceph-disk, users have to migrate sooner or later if they don't want to do everything manually on the CLI (by that I mean via fdisk/pure lvm commands and so on).
We could argue now that this can still be done on the command line manually but all our efforts are towards simplicity/automation and having everything in the Dashboard. If the underlying tool/functionality isn't there anymore, that isn't possible.
I understand your position, yes when we start separating block/db/wal things get really complex that's why I'm sticking with block/db/wal in the same block. Also, we haven't seen any request for separating those when running OSDs on PVC in the Cloud. So we would likely continue to do so for a while.
Also, I'm not saying we should replace the tool but allow not using LVM for a simple scenario to start with
Which then leads me to, why couldn't such functionality be implemented into a single tool instead of having two at the end?
So don't get me wrong, I'm not saying that I'm against everything I'm just saying that I think this is a topic that should be discussed in more depth.
Yes, that's for sure.
As said, just my two cents here.
Kai
-- SUSE Software Solutions Germany GmbH, Maxfeldstr. 5, D 90409 Nürnberg GF:Geschäftsführer: Felix Imendörffer, (HRB 36809, AG Nürnberg)
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
-- Jan Fajerski Senior Software Engineer Enterprise Storage SUSE Software Solutions Germany GmbH Maxfeldstr. 5, 90409 Nürnberg, Germany (HRB 36809, AG Nürnberg) Geschäftsführer: Felix Imendörffer
My two cents, likely controvesial but it has to be said:
So maybe a bare-bones bluestore mode makes sense. In the simple case, it really should be *very* simple. But its scope pretty quickly expodes: what about wal and db devices? We have labels for those, so we could support those, also easily... if the user has to partition the devices beforehand manually.
How is foisting volume management off on the admin any different from foisting off partition management? Today we’re how many years after BlueStore was released? And we *still’* don’t have documentation for planning and managing it that is either complete or correct? Yes 4%, I’m looking at *you*. And don’t get me started about the rbd-mirror ambiguity. We already have bare-bones BlueStore, I’m running thousands of Luminous OSDs on it today. Re WAL and DB, their management and journals before them has indeed been a pain. But those who have to do it have already written tools that work for their scenarios. Moving the target yet again will just mean more time that people have to sink into reinventing the wheel, vs doing interesting things. Or updating their clusters. Changes for the sake of changes have these results: * Time taken away from our burgeoning backlogs to retool for the change du jour * New releases don’t get installed because of the tooling and testing work, and fear of what bizarre pivot-related bug will next break thousands of paying users? There are reasons I’m still running 12.2.2. * Existing bugs persist because devs are working on replacing understood things with non-understood things * New bugs appear for the same reason * We often hear “Why is Ceph so slow? Solution X on the same hardware gives us X times the performance”. Perhaps we could do something about that if we weren’t chasing trends. I feel the same way about containers. What does that bandwagon *really* gain us other than trendiness? How much would they *cost* countless admins who have to throw away years of accumulated know-how and tooling? Too often it seems that Ceph attends to whatever’s whizzy for greenfield deployments, at the expense of jerking around existing, production clusters.
I don't think keeping a simple or barebones approach will survive contact with real-world deployments. Imho if we want a raw mode, we better be prepared to deal with multi-device OSDs and multi-OSD devices and the partitioning this requires.
Real-world deployments have been doing these things for a long time. It is, if not a perfectly solved problem, at least a familiar one. The devil you know.
Since we can't cover all of that, and at a minimum, we can't cover dm-crypt,
dm-crypt is another thing that real-world deployments already have covered. Don’t fix what aint’ broken. That said, dm-crypt is itself a band-aid for hardware shortcomings. SED sure wasn’t manageable. Maybe OPAL will be, if — unlike SMART — manufacturers can surprise us all and deliver usable and interoperable behavior. — aad
On Mon, 9 Dec 2019, Anthony D'Atri wrote:
My two cents, likely controvesial but it has to be said:
[...]
I just want to reiterate, in case it wasn't clear before, that the raw bluestore mode is in no way intended to replace LVM. It explicitly leaves dm-crypt and (for the moment) db/wal out of scope in order to address a narrow use-case (single-volume unencrypted bluestore without lvm). sage
participants (8)
-
Alfredo Deza
-
Anthony D'Atri
-
Blaine Gardner
-
Jan Fajerski
-
Kai Wagner
-
Lars Marowsky-Bree
-
Sage Weil
-
Sebastien Han