Need feedback for Ceph User Survey 2019
Hi all, We conduct yearly user surveys to better under how our users utilize Ceph. The Ceph Foundation collects the data under the Community Data License agreement [0]; which helps the community make more of an informed decision of where our efforts in the development of future releases should go. Back in August, I asked for the community to help draft the next survey [1]. I'm happy to provide a draft of the user survey for 2019. I'm sending this to the dev list in hopes of getting feedback before sending it to the Ceph users list. The first question I received was using something other than Survey monkey due to it not being available in some regions. I have been using another third-party service for our Ceph Days CFP forms, and luckily they offer a survey service that isn't blocked. A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together. Any feedback, corrections or ideas? [0] - https://cdla.io/sharing-1-0/ [1] - https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/Q3NCHOJN45DP... [2] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-general.pdf [3] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-Clusters.pdf -- Mike Perez he/him Ceph Community Manager M: +1-951-572-2633 494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
I feel like some of the options should be multiple choice for people running multiple clusters. I'm assuming that the circles mean a single choice and squares mean multiple choice in your PDFs. 1. The question on which Ceph release do you run should be: "Which Ceph release(s) do you run?" 2. The question on how many OSDs do you have in each node could vary. This could be because a newer sku was added to the cluster which supports more OSDs, or the SSD nodes in a cluster could have a different number of OSDs then the HDD nodes. 3. The OSD backends question could be both BlueStore and FileStore XFS. 4. The types of storage devices should have an Other field (Optane?) Bryan On Oct 1, 2019, at 2:08 PM, Mike Perez <miperez@redhat.com<mailto:miperez@redhat.com>> wrote: Notice: This email is from an external sender. Hi all, We conduct yearly user surveys to better under how our users utilize Ceph. The Ceph Foundation collects the data under the Community Data License agreement [0]; which helps the community make more of an informed decision of where our efforts in the development of future releases should go. Back in August, I asked for the community to help draft the next survey [1]. I'm happy to provide a draft of the user survey for 2019. I'm sending this to the dev list in hopes of getting feedback before sending it to the Ceph users list. The first question I received was using something other than Survey monkey due to it not being available in some regions. I have been using another third-party service for our Ceph Days CFP forms, and luckily they offer a survey service that isn't blocked. A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together. Any feedback, corrections or ideas? [0] - https://cdla.io/sharing-1-0/ [1] - https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/Q3NCHOJN45DP... [2] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-general.pdf [3] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-Clusters.pdf -- Mike Perez he/him Ceph Community Manager M: +1-951-572-2633<tel:+1-951-572-2633> 494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee<https://twitter.com/thingee> Thingee<https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> [https://static.redhat.com/libs/redhat/brand-assets/2/corp/logo--200.png]<https://www.redhat.com/> _______________________________________________ Dev mailing list -- dev@ceph.io<mailto:dev@ceph.io> To unsubscribe send an email to dev-leave@ceph.io<mailto:dev-leave@ceph.io>
Hi Mike Perez, What's about the question "Which type of messenger is configured in the cluster? Posix or RDMA or DPDK" --Thanks Changcheng From: Bryan Stillwell [mailto:bstillwell@godaddy.com] Sent: Wednesday, October 2, 2019 7:03 AM To: Mike Perez <miperez@redhat.com> Cc: dev@ceph.io; Lars Marowsky-Bree <lmb@suse.com> Subject: Re: Need feedback for Ceph User Survey 2019 I feel like some of the options should be multiple choice for people running multiple clusters. I'm assuming that the circles mean a single choice and squares mean multiple choice in your PDFs. 1. The question on which Ceph release do you run should be: "Which Ceph release(s) do you run?" 2. The question on how many OSDs do you have in each node could vary. This could be because a newer sku was added to the cluster which supports more OSDs, or the SSD nodes in a cluster could have a different number of OSDs then the HDD nodes. 3. The OSD backends question could be both BlueStore and FileStore XFS. 4. The types of storage devices should have an Other field (Optane?) Bryan On Oct 1, 2019, at 2:08 PM, Mike Perez <miperez@redhat.com<mailto:miperez@redhat.com>> wrote: Notice: This email is from an external sender. Hi all, We conduct yearly user surveys to better under how our users utilize Ceph. The Ceph Foundation collects the data under the Community Data License agreement [0]; which helps the community make more of an informed decision of where our efforts in the development of future releases should go. Back in August, I asked for the community to help draft the next survey [1]. I'm happy to provide a draft of the user survey for 2019. I'm sending this to the dev list in hopes of getting feedback before sending it to the Ceph users list. The first question I received was using something other than Survey monkey due to it not being available in some regions. I have been using another third-party service for our Ceph Days CFP forms, and luckily they offer a survey service that isn't blocked. A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together. Any feedback, corrections or ideas? [0] - https://cdla.io/sharing-1-0/ [1] - https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/Q3NCHOJN45DP... [2] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-general.pdf [3] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-Clusters.pdf -- Mike Perez he/him Ceph Community Manager M: +1-951-572-2633<tel:+1-951-572-2633> 494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee<https://twitter.com/thingee> Thingee<https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> [https://static.redhat.com/libs/redhat/brand-assets/2/corp/logo--200.png]<https://www.redhat.com/> _______________________________________________ Dev mailing list -- dev@ceph.io<mailto:dev@ceph.io> To unsubscribe send an email to dev-leave@ceph.io<mailto:dev-leave@ceph.io>
On 2019-10-01T16:08:43, Mike Perez <miperez@redhat.com> wrote: Hi all, yay, survey time!
We conduct yearly user surveys to better under how our users utilize Ceph. The Ceph Foundation collects the data under the Community Data License agreement [0]; which helps the community make more of an informed decision of where our efforts in the development of future releases should go.
I also like we're collecting this under the CDLA 1.0 Sharing variant (means we need to avoid any e-mail addresses and org names though, I think; folks probably don't want those globally shared).
A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together.
I think, perhaps, we can just ask for the aggregates across all clusters. If we make it too detailed, it'll be too complex for respondents to enter and we'll not hear from them. Alternatively, we could perhaps have a table for the per-cluster questions, each row representing one cluster or even pool. (e.g., if they have a 10 PiB cluster hosting both hosting 1 PiB replicated metadata and hot data and 6 PiB 8+4 S3 data, what is the answer?) I guess my point is - unless we're going to that level of detail, we're already aggregating (and losing details) at the per-cluster level and per-org aggregation isn't too bad. I'd rather lower the barrier to respond -> "check all that apply" (would make most of our questions multiple choice). We should instead focus on getting that level of detail from Telemetry in 2020. An intermediate solution could be to ask them to "run this command on each of your clusters and paste the output here". But if we ask them to spend 5-10 minutes answering questions per cluster ... (Alas we can't ask them to just turn on Telemetry, not backported and feature complete to all relevant releases. Perhaps we could build a standalone Telemetry client for pre-Nautilus releases? But not for this cycle.) For the fields where we do ask numbers, instead of endless drop-down lists, I'd rather ask for a, well, number. Why give them a drop-down list for total raw capacity? Why not just ask for the number of clusters? How many nodes? Etc, even for replication size/EC profiles. And we should filter redundant questions - if we ask, say, for both the total number of nodes, and the total number of OSDs, we don't have to ask "how many OSDs per node". Unless we consider this a consistency check. The survey pad has a lot of feedback, some of it contradictory, and not all questions asked consistently. So it's not perfectly clear to me what the current consolidated draft would look like. Perhaps if we do that prior to posting to ceph-users, that'd be helpful. Regards, Lars -- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
Hi,
A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together.
Thanks for putting these out for review. I have perhaps written too much in response... :) I think, in general, many of the questions need an "Other" option - e.g. the "Why use Ceph" question In the general questionnaire, you ask about telemetry on "your" cluster, but people might e.g. have several clusters not all of which have telemetry available; similarly, I'd make the "why not" question allow >1 answer - it might be "Well, I run Luminous now, but I don't expect to run telemetry when we move to Nautilus because..." The clusters questionnaire is tricky - it's a lot to fill out if you want folk to do it for every one of their clusters; might you perhaps instead request people fill it out for 1 cluster (production / largest / ...)? It's not clear how I would represent multiple different clusters in the answers to this questionnaire, to be honest, unless they were all essentially the same hardware/software/layout/etc. And I hope that the UI allows e.g. someone who didn't tick that they ran CephFS not to be shown the CephFS questions later. I agree with the suggestion of just letting people type in a number for all the "number of.." questions; if you do that, do you need the "how many OSDs per node" question? "How many OSDs in you cluster?" is a typo (you meant "your") I think. The "number of hours" questions need clarification - do you mean per week? [and "Number of hours troubleshooting the ceph client configuration and software" is duplicated]. I'm guessing these things aren't evenly spread over time, either (e.g. I spend much more time on config/software around a major release bump) I suspect that the "how long to apply updates" question might need a "depends on my vendor" answer? It'd take something like a serious CVE for us to upgrade ourselves rather than waiting for Canonical to produce revised packages, for example, and that can take a while. The answers to the minor updates "why?" don't quite make sense, maybe they've been truncated (e.g. "Concerns about") "What Ceph Manager modules did you enable?" needs a "none" option? [ISTR survey UIs require you to tick a box on a compulsory question] The EC questions often look to be single-answer, but presumably people might run more than one EC setup in the same cluster (for different pools)? It might be worth being clearer about what storage devices you're asking about? (e.g. for OSD storage, for journal/block.db, for host OS, for MON store...) The RBD section should all be multi-answer (except "are you using snapshorts"), I think You're inconsistent as to whether the "snapshots" question wants a reason for either answer (RBD) or for only the No answer (RGW(?),CephFS). Client-side S3, I wonder if s3cmd and minio (Go) might also be worth asking about? Regards, Matthew -- The Wellcome Sanger Institute is operated by Genome Research Limited, a charity registered in England with number 1021457 and a company registered in England with number 2742969, whose registered office is 215 Euston Road, London, NW1 2BE.
I did another pass and I think we can simplify. I suggest we give up on the detailed per-cluster stats from a manual survey and instead rely on telemetry (and/or telemetry backport to mimic/luminous if we *really* want that data), and then consolidate this into a handful of easy questions in the general survey. On the general part, - telemetry 'why' question: make it a "check all that apply" + comment ...the add: - How many clusters do you operate? [fill in number] - Which Ceph releases do you run? Check all that apply (list nautilus -> argonaut) - Total aggregate cluster capacity in TB [fill in number] - Largest cluster capacity in TB [fill in number] and then copy most of the other cluster answers back over to the main one, converting anything that is a selection to a 'check all that a pply'. e.g., - Which Ceph packages (copy/move from cluster survey, but check all that apply) - What operating system(s) are you using on cluster nodes? (copy/move, but check all that apply) A few things could be dropped to simplify: - how many hosts, osds, osds per node, osd backends, redundancy - number of hours (unless we can simplify this?) - data protection scheme - size/min_size - which osd layout features - what type of NICs - RGW: 'do you use snapshots?' subquestion (not a thing) Change: - What process architectures - add Power - Typical number of fs clients (per cluster) - number of files ... (for largest cluster, if multiple clusters) - MDS cache size ... (for largest cluster, if multiple clusters) - number of active MDS for largest cluster [just type in value, not a multiple choice] A few of these still fall into the category of things we should capture with telemetry.. I'm not quite sure where to draw the line, but generally think we should lean toward simplicity. Like, all of those Change items :) Lars, WDYT? sage On Tue, 1 Oct 2019, Mike Perez wrote:
Hi all,
We conduct yearly user surveys to better under how our users utilize Ceph. The Ceph Foundation collects the data under the Community Data License agreement [0]; which helps the community make more of an informed decision of where our efforts in the development of future releases should go.
Back in August, I asked for the community to help draft the next survey [1]. I'm happy to provide a draft of the user survey for 2019. I'm sending this to the dev list in hopes of getting feedback before sending it to the Ceph users list.
The first question I received was using something other than Survey monkey due to it not being available in some regions. I have been using another third-party service for our Ceph Days CFP forms, and luckily they offer a survey service that isn't blocked.
A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together.
Any feedback, corrections or ideas?
[0] - https://cdla.io/sharing-1-0/ [1] - https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/Q3NCHOJN45DP... [2] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-general.pdf [3] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-Clusters.pdf
--
Mike Perez
he/him
Ceph Community Manager
M: +1-951-572-2633
494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
On 2019-10-28T20:32:17, Sage Weil <sage@newdream.net> wrote:
I did another pass and I think we can simplify. I suggest we give up on the detailed per-cluster stats from a manual survey and instead rely on telemetry (and/or telemetry backport to mimic/luminous if we *really* want that data), and then consolidate this into a handful of easy questions in the general survey.
Ack.
On the general part,
- telemetry 'why' question: make it a "check all that apply" + comment
Yes.
...the add:
- How many clusters do you operate? [fill in number] - Which Ceph releases do you run? Check all that apply (list nautilus -> argonaut) - Total aggregate cluster capacity in TB [fill in number] - Largest cluster capacity in TB [fill in number]
Agreed. Though I'd also like to ask about node and OSD counts (aggregate/largest). We could simplify this by collapsing total/largest if they only have one cluster - perhaps we can only show the "largest" column if >=2? The capacity/node/disk count figures always look great in the survey results, and will also allow us to have some comparison relative to the last year's.
- how many hosts, osds, osds per node, osd backends, redundancy - data protection scheme
I'd still like to ask about the data protection schemes, perhaps for the one with the most data? Alternatively, we could ask them for net vs gross vs used capacity and try to compute that from that, but I'm guessing it's easier to ask them for the top 3 schemes they use. The OSD FileStore/BlueStore trends would also be quite useful at this stage, again mainly for comparison purposes. I realize we'll eventually get that from telemetry, but we're not quite there yet.
A few of these still fall into the category of things we should capture with telemetry.. I'm not quite sure where to draw the line, but generally think we should lean toward simplicity. Like, all of those Change items :)
Yes.
Lars, WDYT?
Do we still have a chance to run this survey in 2019? End-of-year survey? Regards, Lars -- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
Here's another draft with the current feedback. Sage, do you have thoughts on keeping any of the data protection questions per Lar's suggestions? On Mon, Oct 28, 2019 at 1:32 PM Sage Weil <sage@newdream.net> wrote:
I did another pass and I think we can simplify. I suggest we give up on the detailed per-cluster stats from a manual survey and instead rely on telemetry (and/or telemetry backport to mimic/luminous if we *really* want that data), and then consolidate this into a handful of easy questions in the general survey.
On the general part,
- telemetry 'why' question: make it a "check all that apply" + comment
...the add:
- How many clusters do you operate? [fill in number] - Which Ceph releases do you run? Check all that apply (list nautilus -> argonaut) - Total aggregate cluster capacity in TB [fill in number] - Largest cluster capacity in TB [fill in number]
and then copy most of the other cluster answers back over to the main one, converting anything that is a selection to a 'check all that a pply'. e.g.,
- Which Ceph packages (copy/move from cluster survey, but check all that apply) - What operating system(s) are you using on cluster nodes? (copy/move, but check all that apply)
A few things could be dropped to simplify:
- how many hosts, osds, osds per node, osd backends, redundancy - number of hours (unless we can simplify this?) - data protection scheme - size/min_size - which osd layout features - what type of NICs - RGW: 'do you use snapshots?' subquestion (not a thing)
Change: - What process architectures - add Power - Typical number of fs clients (per cluster) - number of files ... (for largest cluster, if multiple clusters) - MDS cache size ... (for largest cluster, if multiple clusters) - number of active MDS for largest cluster [just type in value, not a multiple choice]
A few of these still fall into the category of things we should capture with telemetry.. I'm not quite sure where to draw the line, but generally think we should lean toward simplicity. Like, all of those Change items :)
Lars, WDYT?
sage
On Tue, 1 Oct 2019, Mike Perez wrote:
Hi all,
We conduct yearly user surveys to better under how our users utilize Ceph. The Ceph Foundation collects the data under the Community Data License agreement [0]; which helps the community make more of an informed decision of where our efforts in the development of future releases should go.
Back in August, I asked for the community to help draft the next survey [1]. I'm happy to provide a draft of the user survey for 2019. I'm sending this to the dev list in hopes of getting feedback before sending it to the Ceph users list.
The first question I received was using something other than Survey monkey due to it not being available in some regions. I have been using another third-party service for our Ceph Days CFP forms, and luckily they offer a survey service that isn't blocked.
A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together.
Any feedback, corrections or ideas?
[0] - https://cdla.io/sharing-1-0/ [1] -
https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/Q3NCHOJN45DP...
[2] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-general.pdf [3] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-Clusters.pdf
--
Mike Perez
he/him
Ceph Community Manager
M: +1-951-572-2633
494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
-- Mike Perez he/him Ceph Community Manager M: +1-951-572-2633 494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
Hi Mike Perez, Is it possible to add one item to get feedback about which messenger type is used in deployed cluster? TCP-Posix or DPDK or RDMA? B.R. Changcheng From: Mike Perez [mailto:miperez@redhat.com] Sent: Tuesday, November 5, 2019 7:11 AM To: Sage Weil <sage@newdream.net> Cc: dev@ceph.io; Lars Marowsky-Bree <lmb@suse.com> Subject: Re: Need feedback for Ceph User Survey 2019 Here's another draft with the current feedback. Sage, do you have thoughts on keeping any of the data protection questions per Lar's suggestions? On Mon, Oct 28, 2019 at 1:32 PM Sage Weil <sage@newdream.net<mailto:sage@newdream.net>> wrote: I did another pass and I think we can simplify. I suggest we give up on the detailed per-cluster stats from a manual survey and instead rely on telemetry (and/or telemetry backport to mimic/luminous if we *really* want that data), and then consolidate this into a handful of easy questions in the general survey. On the general part, - telemetry 'why' question: make it a "check all that apply" + comment ...the add: - How many clusters do you operate? [fill in number] - Which Ceph releases do you run? Check all that apply (list nautilus -> argonaut) - Total aggregate cluster capacity in TB [fill in number] - Largest cluster capacity in TB [fill in number] and then copy most of the other cluster answers back over to the main one, converting anything that is a selection to a 'check all that a pply'. e.g., - Which Ceph packages (copy/move from cluster survey, but check all that apply) - What operating system(s) are you using on cluster nodes? (copy/move, but check all that apply) A few things could be dropped to simplify: - how many hosts, osds, osds per node, osd backends, redundancy - number of hours (unless we can simplify this?) - data protection scheme - size/min_size - which osd layout features - what type of NICs - RGW: 'do you use snapshots?' subquestion (not a thing) Change: - What process architectures - add Power - Typical number of fs clients (per cluster) - number of files ... (for largest cluster, if multiple clusters) - MDS cache size ... (for largest cluster, if multiple clusters) - number of active MDS for largest cluster [just type in value, not a multiple choice] A few of these still fall into the category of things we should capture with telemetry.. I'm not quite sure where to draw the line, but generally think we should lean toward simplicity. Like, all of those Change items :) Lars, WDYT? sage On Tue, 1 Oct 2019, Mike Perez wrote:
Hi all,
We conduct yearly user surveys to better under how our users utilize Ceph. The Ceph Foundation collects the data under the Community Data License agreement [0]; which helps the community make more of an informed decision of where our efforts in the development of future releases should go.
Back in August, I asked for the community to help draft the next survey [1]. I'm happy to provide a draft of the user survey for 2019. I'm sending this to the dev list in hopes of getting feedback before sending it to the Ceph users list.
The first question I received was using something other than Survey monkey due to it not being available in some regions. I have been using another third-party service for our Ceph Days CFP forms, and luckily they offer a survey service that isn't blocked.
A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together.
Any feedback, corrections or ideas?
[0] - https://cdla.io/sharing-1-0/ [1] - https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/Q3NCHOJN45DP... [2] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-general.pdf [3] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-Clusters.pdf
--
Mike Perez
he/him
Ceph Community Manager
M: +1-951-572-2633
494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
-- Mike Perez he/him Ceph Community Manager M: +1-951-572-2633<tel:+1-951-572-2633> 494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee<https://twitter.com/thingee> Thingee<https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> [https://static.redhat.com/libs/redhat/brand-assets/2/corp/logo--200.png]<https://www.redhat.com>
On Mon, 4 Nov 2019, Mike Perez wrote:
Here's another draft with the current feedback. Sage, do you have thoughts on keeping any of the data protection questions per Lar's suggestions?
No opinion... so sure. Few small edits: - Total raw and usable capacity: let them type in a box (TB) - Which ceph release question should be release(s), w/ checkboxes - What Ceph Manager modules question: s/did/do/ - RGW auth question: checkboxes, not radio buttons - number of rgw sites: "Number of federated RGW sites", with an entry box Then let's ship it! Thanks! sage
On Mon, Oct 28, 2019 at 1:32 PM Sage Weil <sage@newdream.net> wrote:
I did another pass and I think we can simplify. I suggest we give up on the detailed per-cluster stats from a manual survey and instead rely on telemetry (and/or telemetry backport to mimic/luminous if we *really* want that data), and then consolidate this into a handful of easy questions in the general survey.
On the general part,
- telemetry 'why' question: make it a "check all that apply" + comment
...the add:
- How many clusters do you operate? [fill in number] - Which Ceph releases do you run? Check all that apply (list nautilus -> argonaut) - Total aggregate cluster capacity in TB [fill in number] - Largest cluster capacity in TB [fill in number]
and then copy most of the other cluster answers back over to the main one, converting anything that is a selection to a 'check all that a pply'. e.g.,
- Which Ceph packages (copy/move from cluster survey, but check all that apply) - What operating system(s) are you using on cluster nodes? (copy/move, but check all that apply)
A few things could be dropped to simplify:
- how many hosts, osds, osds per node, osd backends, redundancy - number of hours (unless we can simplify this?) - data protection scheme - size/min_size - which osd layout features - what type of NICs - RGW: 'do you use snapshots?' subquestion (not a thing)
Change: - What process architectures - add Power - Typical number of fs clients (per cluster) - number of files ... (for largest cluster, if multiple clusters) - MDS cache size ... (for largest cluster, if multiple clusters) - number of active MDS for largest cluster [just type in value, not a multiple choice]
A few of these still fall into the category of things we should capture with telemetry.. I'm not quite sure where to draw the line, but generally think we should lean toward simplicity. Like, all of those Change items :)
Lars, WDYT?
sage
On Tue, 1 Oct 2019, Mike Perez wrote:
Hi all,
We conduct yearly user surveys to better under how our users utilize Ceph. The Ceph Foundation collects the data under the Community Data License agreement [0]; which helps the community make more of an informed decision of where our efforts in the development of future releases should go.
Back in August, I asked for the community to help draft the next survey [1]. I'm happy to provide a draft of the user survey for 2019. I'm sending this to the dev list in hopes of getting feedback before sending it to the Ceph users list.
The first question I received was using something other than Survey monkey due to it not being available in some regions. I have been using another third-party service for our Ceph Days CFP forms, and luckily they offer a survey service that isn't blocked.
A second question that came up was how to layout questions for multiple cluster deployments. An idea I had was having our general Ceph user survey [2] separate from the deployment questions [3]. The general questions only need to be answered once, and the deployment survey can be answered multiple times to capture the different configurations. I'm looking into a way to link the answers of both surveys together.
Any feedback, corrections or ideas?
[0] - https://cdla.io/sharing-1-0/ [1] -
https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/Q3NCHOJN45DP...
[2] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-general.pdf [3] - https://ceph.io/wp-content/uploads/2019/10/Ceph-User-Survey-Clusters.pdf
--
Mike Perez
he/him
Ceph Community Manager
M: +1-951-572-2633
494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
--
Mike Perez
he/him
Ceph Community Manager
M: +1-951-572-2633
494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
On Tue, Nov 5, 2019 at 12:11 AM Mike Perez <miperez@redhat.com> wrote:
Here's another draft with the current feedback. Sage, do you have thoughts on keeping any of the data protection questions per Lar's suggestions?
- Do you using cache tiering? s/using/use - In "How soon do you apply dot/minor releases to your cluster? Why?" I see "Concerns about" and "Concerns about new". Was that supposed to be "Concerns about stability", etc? - RBD section is the only one that has "Environment status" question. Should we add it to RGW and CephFS sections as well? Thanks, Ilya
On 2019-11-04T15:11:11, Mike Perez <miperez@redhat.com> wrote:
Here's another draft with the current feedback. Sage, do you have thoughts on keeping any of the data protection questions per Lar's suggestions?
In addition to the comments from others: In the CephFS section, we ask about the "typical" number of fs clients. Everywhere else we ask about the aggregate. Is that intentional? I'm just wondering if this tells us anything, e.g., imagine someone responding with 5 clusters, only one of them with CephFS, "typical" then would be "1-5"? Should perhaps ask about the largest, like the next question - or the total? For active MDS, we again ask for the total and also the largest number. Should we do the same for the previous questions for consistency? Again, I'd turn those into number entry fields rather than drop down. Snapshots -> No, please specify "why not" Everything else looks good. Joao also wants to look into if the telemetry plugin may be backportable to Luminous. (I don't think earlier makes sense.) That'll help with the data set for the slow adopters. Regards, Lars -- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
Another pass on the survey with everyone's suggestions. On Tue, Nov 5, 2019 at 6:44 AM Lars Marowsky-Bree <lmb@suse.com> wrote:
On 2019-11-04T15:11:11, Mike Perez <miperez@redhat.com> wrote:
Here's another draft with the current feedback. Sage, do you have thoughts on keeping any of the data protection questions per Lar's suggestions?
In addition to the comments from others:
In the CephFS section, we ask about the "typical" number of fs clients. Everywhere else we ask about the aggregate. Is that intentional?
I'm just wondering if this tells us anything, e.g., imagine someone responding with 5 clusters, only one of them with CephFS, "typical" then would be "1-5"?
Should perhaps ask about the largest, like the next question - or the total?
For active MDS, we again ask for the total and also the largest number. Should we do the same for the previous questions for consistency?
Again, I'd turn those into number entry fields rather than drop down.
Snapshots -> No, please specify "why not"
Everything else looks good.
Joao also wants to look into if the telemetry plugin may be backportable to Luminous. (I don't think earlier makes sense.) That'll help with the data set for the slow adopters.
Regards, Lars
-- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
-- Mike Perez he/him Ceph Community Manager M: +1-951-572-2633 494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
On 2019-11-08T11:48:39, Mike Perez <miperez@redhat.com> wrote:
Another pass on the survey with everyone's suggestions.
Thanks! Much appreciated, and I think we could roll this. Though the sizing questions still leave me confused. Why can't they enter numbers for their total raw capacity? For CephFS: I just noticed that for the clients it's ranges, for files#, it's an order of magnitude (I have 5mil, what do I choose? And what does 10+ mean?), then ranges (two of them overlapping) for MDS cache size, and number of MDS is actually asked twice (with very narrow ranges, and then as free entry; remove the former?). I think just replacing these with number entry fields would be best. But I'm not the CephFS expert who asked them, so take that with a large pinch of salt. Let's roll and then make sure everyone gets telemetry enabled so we can stop asking :-D Regards, Lars -- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
On 2019-11-08T11:48:39, Mike Perez <miperez@redhat.com> wrote: Sage just shared the link to the survey on Zoho. Thanks! Minor feedback (sorry about that, some things only become clear seeing the actual forms) that'd be nice to address but are optional: Page 1: - "Is telemetry enabled in your cluster*s*" (might have more than one) And then also add "Partially" as an answer. - Can we had the "If not" sections unless folks have not picked "Yes"? - "How do you participate in the Ceph community?" - Add "Member of Ceph Foundation" as an answer ;-) - The country list is intimidating; is there any way with Zoho to make this a drop down selection or tag cloud? (This is purely visual) Page 2: - Total usable capacity should also be a number entry to match the previous questions - Is someone possibly running "Master" or do we rather not even dare ask? - "Which Ceph Manager modules" probably should have a "I do not know" answer, or "Not applicable" (if < Luminous), and definitely also "Telemetry" / "Prometheus" - Messenger type (TCP-Posix) deserves being the default, or adding a (Default) note in case users are confused Page 3: - IHV -> add "Prefer not to say" Page 4: - "What are the use cases" -> add "Backup" Page RBD: - "Do you use snapshots" -> No already opens a text entry field, the last "Why?" is perhaps redundant on that page? Page RGW: - That the workload choices differ across all pages is highly annoying and will make data analysis harder ;-) Page CephFS: - See previous comment - Everywhere else the environment status is before the workload question - My comments regarding "Number of files" being unclear about which way to round have not been addressed. I'd suggest to add a "< " in front of the fields instead of a "~" to turn them into ranges. Thanks!, Lars -- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
Hi, => How about adding question about running on bare metal vs. virtualized ? => How about asking for layout options, like single, stretched, how many DCs stretched over ? => perhaps, too, how many OSDs in the biggest cluster ? H/W section: in "Which type of storage devices are used?" => could add Optane => and could add specific question for NVMe and Optane for the use case (e.g., RocksDB/WAL, OSD, OSD for object index) in "Which OSD layout features do you use?" => s/Seperate/Separate/ journal device (FileStore) => could add separation of "RocksDB and WAL and data" RGW section: => How about asking - for the number of RGW ? - for largest objects stored ? - how many objects in buckets, how many buckets ? - how many realms, max # zonegroups per realm, max # of zones per zonegroup ? - how about asking RGW multi-site more detailed ? (as max values, how many sites, how many zonegroups, how many zones per group) (we ask more detailed in MDS section, too) CephFS section: in "Number of files (getfattr -d -m ceph.dir.rfiles /mnt/cephfs) (for largest cluster, if multiple clusters)": the last should be something 10G+ or better use for M and G the more common 'm' and 'bn' for numbers. Don't we wanna ask for some wish list features/functionality with prio ? (So, may be 2-3 items max.) G, -matt On 22.11.19 15:51, Lars Marowsky-Bree wrote:
On 2019-11-08T11:48:39, Mike Perez <miperez@redhat.com> wrote:
Sage just shared the link to the survey on Zoho. Thanks!
Minor feedback (sorry about that, some things only become clear seeing the actual forms) that'd be nice to address but are optional:
Page 1: - "Is telemetry enabled in your cluster*s*" (might have more than one) And then also add "Partially" as an answer.
- Can we had the "If not" sections unless folks have not picked "Yes"?
- "How do you participate in the Ceph community?" - Add "Member of Ceph Foundation" as an answer ;-)
- The country list is intimidating; is there any way with Zoho to make this a drop down selection or tag cloud? (This is purely visual)
Page 2: - Total usable capacity should also be a number entry to match the previous questions
- Is someone possibly running "Master" or do we rather not even dare ask?
- "Which Ceph Manager modules" probably should have a "I do not know" answer, or "Not applicable" (if < Luminous), and definitely also "Telemetry" / "Prometheus"
- Messenger type (TCP-Posix) deserves being the default, or adding a (Default) note in case users are confused
Page 3: - IHV -> add "Prefer not to say"
Page 4: - "What are the use cases" -> add "Backup"
Page RBD: - "Do you use snapshots" -> No already opens a text entry field, the last "Why?" is perhaps redundant on that page?
Page RGW: - That the workload choices differ across all pages is highly annoying and will make data analysis harder ;-)
Page CephFS: - See previous comment
- Everywhere else the environment status is before the workload question
- My comments regarding "Number of files" being unclear about which way to round have not been addressed. I'd suggest to add a "< " in front of the fields instead of a "~" to turn them into ranges.
Thanks!, Lars
-- —————————————————— Matthias Muench Senior Specialist Solution Architect EMEA Storage Specialist matthias.muench@redhat.com Phone: +49-160-92654111 Red Hat GmbH Werner-von-Siemens-Ring 14 85630 Grasbrunn Germany _______________________________________________________________________ Red Hat GmbH, http://www.de.redhat.com · Registered seat: Grasbrunn, Commercial register: Amtsgericht Muenchen HRB 153243 · Managing Directors: Charles Cachera, Michael O'Neill, Tom Savage, Eric Shander
Replies inline... On Fri, Nov 22, 2019 at 6:51 AM Lars Marowsky-Bree <lmb@suse.com> wrote:
On 2019-11-08T11:48:39, Mike Perez <miperez@redhat.com> wrote:
Sage just shared the link to the survey on Zoho. Thanks!
Minor feedback (sorry about that, some things only become clear seeing the actual forms) that'd be nice to address but are optional:
Page 1: - "Is telemetry enabled in your cluster*s*" (might have more than one) And then also add "Partially" as an answer.
I'm not sure if it would be useful to capture multiple answers here. Can you elaborate? - Can we had the "If not" sections unless folks have not picked "Yes"?
Done
- "How do you participate in the Ceph community?" - Add "Member of Ceph Foundation" as an answer ;-)
Done
- The country list is intimidating; is there any way with Zoho to make this a drop down selection or tag cloud? (This is purely visual)
I was not able to find another input option. We could do a multiline text field and have the responder type one on each line.
Page 2: - Total usable capacity should also be a number entry to match the previous questions
Done
- Is someone possibly running "Master" or do we rather not even dare ask?
I made it an option in the release question.
- "Which Ceph Manager modules" probably should have a "I do not know" answer, or "Not applicable" (if < Luminous), and definitely also "Telemetry" / "Prometheus"
Done
- Messenger type (TCP-Posix) deserves being the default, or adding a (Default) note in case users are confused
Done
Page 3: - IHV -> add "Prefer not to say"
Done Page 4:
- "What are the use cases" -> add "Backup"
Done
Page RBD: - "Do you use snapshots" -> No already opens a text entry field, the last "Why?" is perhaps redundant on that page?
Done
Page RGW: - That the workload choices differ across all pages is highly annoying and will make data analysis harder ;-)
Done
Page CephFS: - See previous comment
Done
- Everywhere else the environment status is before the workload question
I removed the general one and added a match question/answers to each component page. - My comments regarding "Number of files" being unclear about which way
to round have not been addressed. I'd suggest to add a "< " in front of the fields instead of a "~" to turn them into ranges.
Done -- Mike Perez he/him Ceph Community Manager M: +1-951-572-2633 494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
On 2019-11-25T16:52:53, Mike Perez <miperez@redhat.com> wrote: Perfect, thanks!
Page 1: - "Is telemetry enabled in your cluster*s*" (might have more than one) And then also add "Partially" as an answer. I'm not sure if it would be useful to capture multiple answers here. Can you elaborate?
Well, we are asking about multiple clusters in the rest of the survey, and they may be at different versions, or in different environments. Answering "Yes/No" doesn't fully capture the possible state, and how to proceed in that case is unclear. I'd add it just for completeness. (But looking at the data analysis later, if someone answered "partially", "Not available in version X", and then flagged, say, Luminous/Nautilus, one might think that this validates the effort of backporting.)
- The country list is intimidating; is there any way with Zoho to make this a drop down selection or tag cloud? (This is purely visual) I was not able to find another input option. We could do a multiline text field and have the responder type one on each line.
No, that's fine. I was just wondering, no reason to hold the survey back. This looks good to roll out now from my PoV! Thanks for being so patient with my nitpicking, I really appreciate that. Regards, Lars -- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
Thanks everyone for the feedback. I'm pleased to announce the survey is now live: https://ceph.io/user-survey/ https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/XTIAIA427G75... On Tue, Nov 26, 2019 at 12:36 AM Lars Marowsky-Bree <lmb@suse.com> wrote:
On 2019-11-25T16:52:53, Mike Perez <miperez@redhat.com> wrote:
Perfect, thanks!
Page 1: - "Is telemetry enabled in your cluster*s*" (might have more than one) And then also add "Partially" as an answer. I'm not sure if it would be useful to capture multiple answers here. Can you elaborate?
Well, we are asking about multiple clusters in the rest of the survey, and they may be at different versions, or in different environments. Answering "Yes/No" doesn't fully capture the possible state, and how to proceed in that case is unclear. I'd add it just for completeness.
(But looking at the data analysis later, if someone answered "partially", "Not available in version X", and then flagged, say, Luminous/Nautilus, one might think that this validates the effort of backporting.)
- The country list is intimidating; is there any way with Zoho to make this a drop down selection or tag cloud? (This is purely visual) I was not able to find another input option. We could do a multiline text field and have the responder type one on each line.
No, that's fine. I was just wondering, no reason to hold the survey back.
This looks good to roll out now from my PoV! Thanks for being so patient with my nitpicking, I really appreciate that.
Regards, Lars
-- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden) _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
-- Mike Perez he/him Ceph Community Manager M: +1-951-572-2633 494C 5D25 2968 D361 65FB 3829 94BC D781 ADA8 8AEA @Thingee <https://twitter.com/thingee> Thingee <https://www.linkedin.com/thingee> <https://www.facebook.com/RedHatInc> <https://www.redhat.com>
participants (8)
-
Bryan Stillwell
-
Ilya Dryomov
-
Lars Marowsky-Bree
-
Liu, Changcheng
-
Matthew Vernon
-
Matthias Muench
-
Mike Perez
-
Sage Weil