Hi, Can you suggest me what is a good cephfs design? I've never used it, only rgw and rbd we have, but want to give a try. Howvere in the mail list I saw a huge amount of issues with cephfs so would like to go with some let's say bulletproof best practices. Like separate the mds from mon and mgr? Need a lot of memory? Should be on ssd or nvme? How many cpu/disk ... Very appreciate it. Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com> --------------------------------------------------- ________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses.
Hi, first of all, check the workload you like to have on the filesystem if you plan to migrate an old one do some proper performance-testing of the old storage. the io500 can give some ideas https://www.vi4io.org/io500/start but it depends on the use-case of the filesystem cheers, Ansgar Am Fr., 11. Juni 2021 um 10:54 Uhr schrieb Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com>:
Hi,
Can you suggest me what is a good cephfs design? I've never used it, only rgw and rbd we have, but want to give a try. Howvere in the mail list I saw a huge amount of issues with cephfs so would like to go with some let's say bulletproof best practices.
Like separate the mds from mon and mgr? Need a lot of memory? Should be on ssd or nvme? How many cpu/disk ...
Very appreciate it.
Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com> ---------------------------------------------------
________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Couple of team want to use cephfs with k8s so the main use case would be k8s users. Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com> --------------------------------------------------- On 2021. Jun 11., at 17:48, Ansgar Jazdzewski <a.jazdzewski@googlemail.com> wrote: Hi, first of all, check the workload you like to have on the filesystem if you plan to migrate an old one do some proper performance-testing of the old storage. the io500 can give some ideas https://www.vi4io.org/io500/start but it depends on the use-case of the filesystem cheers, Ansgar Am Fr., 11. Juni 2021 um 10:54 Uhr schrieb Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com>: Hi, Can you suggest me what is a good cephfs design? I've never used it, only rgw and rbd we have, but want to give a try. Howvere in the mail list I saw a huge amount of issues with cephfs so would like to go with some let's say bulletproof best practices. Like separate the mds from mon and mgr? Need a lot of memory? Should be on ssd or nvme? How many cpu/disk ... Very appreciate it. Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com> --------------------------------------------------- ________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
hey Istvan, The Hardware Recommendations <https://docs.ceph.com/en/latest/start/hardware-recommendations/> page actually has a ton of info on the questions you are asking, did you go though that one yet? Without massive overkill, I don't think there's a "bulletproof" design, as the actual I/O use cases vary wildly depending on the application. e.g. you mentioned that "main use case would be k8s users", are there 2000 users that need 500 IOPS each at 100MB/s, OR 5 users that touch the storage every few minutes and store 10MB of data. These are polar opposites with orders of magnitude different requirements. On Fri, Jun 11, 2021 at 10:56 AM Szabo, Istvan (Agoda) < Istvan.Szabo@agoda.com> wrote:
Couple of team want to use cephfs with k8s so the main use case would be k8s users.
Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com> ---------------------------------------------------
On 2021. Jun 11., at 17:48, Ansgar Jazdzewski <a.jazdzewski@googlemail.com> wrote:
Hi,
first of all, check the workload you like to have on the filesystem if you plan to migrate an old one do some proper performance-testing of the old storage.
the io500 can give some ideas https://www.vi4io.org/io500/start but it depends on the use-case of the filesystem
cheers, Ansgar
Am Fr., 11. Juni 2021 um 10:54 Uhr schrieb Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com>:
Hi,
Can you suggest me what is a good cephfs design? I've never used it, only rgw and rbd we have, but want to give a try. Howvere in the mail list I saw a huge amount of issues with cephfs so would like to go with some let's say bulletproof best practices.
Like separate the mds from mon and mgr? Need a lot of memory? Should be on ssd or nvme? How many cpu/disk ...
Very appreciate it.
Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com> ---------------------------------------------------
________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Cheers, Peter Sarossy Technical Program Manager Data Center Data Security - Google LLC.
Can you suggest me what is a good cephfs design?
One that uses copious complements of my employer’s components, naturally ;)
I've never used it, only rgw and rbd we have, but want to give a try. Howvere in the mail list I saw a huge amount of issues with cephfs
Something to remember about the list is that people are far more likely to post when they have a problem than when things are running fine, so it’s easy to mistake that for instabiliity. For every issue posted, there are a bunch of clusters humming right along.
so would like to go with some let's say bulletproof best practices.
Like separate the mds from mon and mgr? Need a lot of memory? Should be on ssd or nvme? How many cpu/disk ...
Like Peter wrote, that’s very dependent on the scale and nature of your workload.
On any given a properly sized ceph setup, for other than database end use) theoretically shouldn't a ceph-fs root out-perform any fs atop a rados block device root? Seems to me like it ought to: moving only the 'interesting' bits of files over the so-called 'public' network should take fewer, smaller packets than the overhead associated with whole blocks that hold some fraction of the 'interesting' bits? Does it work that way in practice?
Not really clear for me to be ho mnest how many cephfs is needed? Would it worth to create multiple or what is the use case to create multiple? In the examples and how people using it seems like only 1 cephfs + 1 metadata pool on nvme, not really multiple cephfs. Doc just relates that if you want multiple cephfs what to do but not too much information about use case. Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com> --------------------------------------------------- On 2021. Jun 11., at 23:31, Harry G. Coin <hgcoin@gmail.com> wrote: On any given a properly sized ceph setup, for other than database end use) theoretically shouldn't a ceph-fs root out-perform any fs atop a rados block device root? Seems to me like it ought to: moving only the 'interesting' bits of files over the so-called 'public' network should take fewer, smaller packets than the overhead associated with whole blocks that hold some fraction of the 'interesting' bits? Does it work that way in practice? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io ________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses.
I doubt it. The problem is that the CephFS MDS must perform distributed metadata transactions with ordering and locking. Whereas a filesystem on rbd runs locally and doesn't have to worry about other computers writing to the same block device. Our bottleneck in production is usually the MDS CPU load. On Fri, Jun 11, 2021 at 12:31 PM Harry G. Coin <hgcoin@gmail.com> wrote:
On any given a properly sized ceph setup, for other than database end use) theoretically shouldn't a ceph-fs root out-perform any fs atop a rados block device root?
Seems to me like it ought to: moving only the 'interesting' bits of files over the so-called 'public' network should take fewer, smaller packets than the overhead associated with whole blocks that hold some fraction of the 'interesting' bits?
Does it work that way in practice?
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On 6/12/21 4:39 PM, Nathan Fish wrote:
I doubt it. The problem is that the CephFS MDS must perform distributed metadata transactions with ordering and locking. Whereas a filesystem on rbd runs locally and doesn't have to worry about other computers writing to the same block device. Our bottleneck in production is usually the MDS CPU load.
Perhaps if an 'exclusive write mount' option existed, the mds could delegate most of what it does to the client. Moving 4K-at-least blocks around a network, even with 'jumbo frames' on local segments, has got to take more processing and network bandwidth than the 'known interesting only' parts of files.
On Fri, Jun 11, 2021 at 12:31 PM Harry G. Coin <hgcoin@gmail.com> wrote:
On any given a properly sized ceph setup, for other than database end use) theoretically shouldn't a ceph-fs root out-perform any fs atop a rados block device root?
Seems to me like it ought to: moving only the 'interesting' bits of files over the so-called 'public' network should take fewer, smaller packets than the overhead associated with whole blocks that hold some fraction of the 'interesting' bits?
Does it work that way in practice?
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Peter, Yeah, went through all, also set the mds_memory_limit. Collocated with mgr/mon so created 3 mds. Have enough cpu, it is on ssd. So yeah, even went inside all the cephfs menu points to take the information. Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com> --------------------------------------------------- On 2021. Jun 11., at 22:36, Peter Sarossy <peter.sarossy@gmail.com> wrote: hey Istvan, The Hardware Recommendations<https://docs.ceph.com/en/latest/start/hardware-recommendations/> page actually has a ton of info on the questions you are asking, did you go though that one yet? Without massive overkill, I don't think there's a "bulletproof" design, as the actual I/O use cases vary wildly depending on the application. e.g. you mentioned that "main use case would be k8s users", are there 2000 users that need 500 IOPS each at 100MB/s, OR 5 users that touch the storage every few minutes and store 10MB of data. These are polar opposites with orders of magnitude different requirements. On Fri, Jun 11, 2021 at 10:56 AM Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com<mailto:Istvan.Szabo@agoda.com>> wrote: Couple of team want to use cephfs with k8s so the main use case would be k8s users. Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com><mailto:istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com>> --------------------------------------------------- On 2021. Jun 11., at 17:48, Ansgar Jazdzewski <a.jazdzewski@googlemail.com<mailto:a.jazdzewski@googlemail.com>> wrote: Hi, first of all, check the workload you like to have on the filesystem if you plan to migrate an old one do some proper performance-testing of the old storage. the io500 can give some ideas https://www.vi4io.org/io500/start but it depends on the use-case of the filesystem cheers, Ansgar Am Fr., 11. Juni 2021 um 10:54 Uhr schrieb Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com<mailto:Istvan.Szabo@agoda.com>>: Hi, Can you suggest me what is a good cephfs design? I've never used it, only rgw and rbd we have, but want to give a try. Howvere in the mail list I saw a huge amount of issues with cephfs so would like to go with some let's say bulletproof best practices. Like separate the mds from mon and mgr? Need a lot of memory? Should be on ssd or nvme? How many cpu/disk ... Very appreciate it. Istvan Szabo Senior Infrastructure Engineer --------------------------------------------------- Agoda Services Co., Ltd. e: istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com><mailto:istvan.szabo@agoda.com<mailto:istvan.szabo@agoda.com>> --------------------------------------------------- ________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io<mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io<mailto:ceph-users-leave@ceph.io> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io<mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io<mailto:ceph-users-leave@ceph.io> -- Cheers, Peter Sarossy Technical Program Manager Data Center Data Security - Google LLC.
On 6/11/21 10:52 AM, Szabo, Istvan (Agoda) wrote:
Hi,
Can you suggest me what is a good cephfs design? I've never used it, only rgw and rbd we have, but want to give a try. Howvere in the mail list I saw a huge amount of issues with cephfs so would like to go with some let's say bulletproof best practices.
You want to use as less features as possible: no snapshotting, no multi MDS, no multi cephfs. Standby-replay is fine though. All these features are nice, and have seen a lot of improvements, but those are also the parts that have become "stable" most recently. A lot of "issues" that are discussed on the list with regards to CephFS are related to above features one way or another. There *used* to be a part in the docs that more or less stated the above. It might still be there, but I can't find it. YMMV. Gr. Stefan
On 6/11/21 3:52 AM, Szabo, Istvan (Agoda) wrote:
Hi,
Can you suggest me what is a good cephfs design? I've never used it, only rgw and rbd we have, but want to give a try. Howvere in the mail list I saw a huge amount of issues with cephfs so would like to go with some let's say bulletproof best practices.
You've read many practical answers to your question so far. My contribution is: cephfs has to 'win' over the long term because moving 'known interesting' data over a network will always take less time than having a client 'file system' move whole storage blocks over the fiber or wire then have to sort out the bits the application actually wants. The only way that doesn't happen is if the 'wires' are dramatically faster than the hosts and lightly loaded -- not what's expected. So, long term: cephfs has the logical ability to out-perform other block-backed (rbd/iscsi) choices. But not today. The thing that makes it 'seem slow' now is dealing with the multi-user file/record level contention block devices don't have to face. Over time I expect directory trees might be shared with a 'one user' flag that might allow the client to interact with the mons/osds directly and require very little mds traffic. That will win over rbd+fs designs because of 'more of what the user wants per network packet' issues. So, eventually (Year? Years? Decades?) I think RadosGW and cephfs will bear most of the ceph traffic. But for today -- if a host is the sole user of a directory tree -- rbd + xfs (ymmv) HC
participants (7)
-
Ansgar Jazdzewski
-
Anthony D'Atri
-
Harry G. Coin
-
Nathan Fish
-
Peter Sarossy
-
Stefan Kooman
-
Szabo, Istvan (Agoda)