Possible performance regression between 15.2.4 and HEAD?
Hi all, There appears to be a performance regression going from 15.2.4 to HEAD. I first realized this when testing my patches to Ceph on an 8-node cluster, but it is easily reproducible on *vanilla* Ceph with vstart as well, using the following steps: $ git clone https://github.com/ceph/ceph.git && cd ceph $ ./do_cmake.sh -DCMAKE_BUILD_TYPE=RelWithDebInfo -DWITH_MANPAGE=OFF -DWITH_BABELTRACE=OFF -DWITH_MGR_DASHBOARD_FRONTEND=OFF && cd build && make -j32 vstart $ MON=1 OSD=1 MDS=0 ../src/vstart.sh --debug --new --localhost --bluestore --bluestore-devs /dev/xxx $ sudo ./bin/ceph osd pool create foo 32 32 $ sudo ./bin/rados bench -p foo 100 write --no-cleanup With the old hard drive that I have (Hitachi HUA72201), I'm getting an average throughput of 60 MiB/s. When I switch to v15.2.4 (git checkout v15.2.4), rebuild, and repeat the experiment, and I get an average throughput of 90 MiB/s. I've reliably reproduced similar difference between 15.2.4 and HEAD by building release packages and running them on an 8-node cluster. Is this expected or is this a performance regression? Thanks!
Hi Abutalib, Given that you are on HDDs, I'd look closely at Igor's bluestore disk allocator work, though I'm not sure how many of his changes are active right now. Especially look closely at the min_alloc_size change. https://github.com/ceph/ceph/pull/33365 https://github.com/ceph/ceph/pull/34588 In the past I've found that the most reliable way to figure out these kind of issues is to do a git bisect and run the tests each time. It's a pain and can take forever but it usually does a pretty good job of pinpointing what's going on (even if it sometimes points out that it was my fault due to a test error!). I'd check out the allocator work first though. Mark On 8/3/20 10:29 PM, Abutalib Aghayev wrote:
Hi all,
There appears to be a performance regression going from 15.2.4 to HEAD. I first realized this when testing my patches to Ceph on an 8-node cluster, but it is easily reproducible on *vanilla* Ceph with vstart as well, using the following steps:
$ git clone https://github.com/ceph/ceph.git && cd ceph $ ./do_cmake.sh -DCMAKE_BUILD_TYPE=RelWithDebInfo -DWITH_MANPAGE=OFF -DWITH_BABELTRACE=OFF -DWITH_MGR_DASHBOARD_FRONTEND=OFF && cd build && make -j32 vstart $ MON=1 OSD=1 MDS=0 ../src/vstart.sh --debug --new --localhost --bluestore --bluestore-devs /dev/xxx $ sudo ./bin/ceph osd pool create foo 32 32 $ sudo ./bin/rados bench -p foo 100 write --no-cleanup
With the old hard drive that I have (Hitachi HUA72201), I'm getting an average throughput of 60 MiB/s. When I switch to v15.2.4 (git checkout v15.2.4), rebuild, and repeat the experiment, and I get an average throughput of 90 MiB/s. I've reliably reproduced similar difference between 15.2.4 and HEAD by building release packages and running them on an 8-node cluster.
Is this expected or is this a performance regression?
Thanks!
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Hi Mark, Given that it is a pain and takes forever, it would be great to automate it, like some other projects already do: https://chromium.googlesource.com/chromium/src/+/master/docs/speed/addressin.... I realize that it is a nontrivial amount of work, but perhaps something to keep in mind for a GSOC project. On Tue, Aug 4, 2020 at 9:29 AM Mark Nelson <mnelson@redhat.com> wrote:
Hi Abutalib,
Given that you are on HDDs, I'd look closely at Igor's bluestore disk allocator work, though I'm not sure how many of his changes are active right now. Especially look closely at the min_alloc_size change.
https://github.com/ceph/ceph/pull/33365
https://github.com/ceph/ceph/pull/34588
In the past I've found that the most reliable way to figure out these kind of issues is to do a git bisect and run the tests each time. It's a pain and can take forever but it usually does a pretty good job of pinpointing what's going on (even if it sometimes points out that it was my fault due to a test error!). I'd check out the allocator work first though.
Mark
On 8/3/20 10:29 PM, Abutalib Aghayev wrote:
Hi all,
There appears to be a performance regression going from 15.2.4 to HEAD. I first realized this when testing my patches to Ceph on an 8-node cluster, but it is easily reproducible on *vanilla* Ceph with vstart as well, using the following steps:
$ git clone https://github.com/ceph/ceph.git && cd ceph $ ./do_cmake.sh -DCMAKE_BUILD_TYPE=RelWithDebInfo -DWITH_MANPAGE=OFF -DWITH_BABELTRACE=OFF -DWITH_MGR_DASHBOARD_FRONTEND=OFF && cd build && make -j32 vstart $ MON=1 OSD=1 MDS=0 ../src/vstart.sh --debug --new --localhost --bluestore --bluestore-devs /dev/xxx $ sudo ./bin/ceph osd pool create foo 32 32 $ sudo ./bin/rados bench -p foo 100 write --no-cleanup
With the old hard drive that I have (Hitachi HUA72201), I'm getting an average throughput of 60 MiB/s. When I switch to v15.2.4 (git checkout v15.2.4), rebuild, and repeat the experiment, and I get an average throughput of 90 MiB/s. I've reliably reproduced similar difference between 15.2.4 and HEAD by building release packages and running them on an 8-node cluster.
Is this expected or is this a performance regression?
Thanks!
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Hi All, I've just completed a small investigation on this issue. Short summary for Abutalib's benchmark scenario in my environment: - when "vstarted" cluster has WAL/DB at NVMe drive - master is ~65% faster (130 MB/s vs. 75MB/s) - in spinner-only setup (WAL/DB are at the same spinner) - octopus is 70 % faster than master ( 88MB/s vs. 52 MB/s) Looks like it is bluestore_max_blob_size_hdd set to 64K by default which causes performance drop in this case (commit sha a8733598eddf57dca86bf002653e630a7cf4db6e). Setting it back to 512K results in a better performance for master: 113 MB/s vs. 88 MB/s But I presume these numbers are tightly bound to writing 4MB chunks benchmark scenario. Different chunk size will most probably cause pretty different results. Hence reverting bluestore_max_blob_size_hdd back to 512K by default is IMO still questionable. Wondering if the above results are in-line with somebody's else ones.... Thanks, Igor On 8/4/2020 6:29 AM, Abutalib Aghayev wrote:
Hi all,
There appears to be a performance regression going from 15.2.4 to HEAD. I first realized this when testing my patches to Ceph on an 8-node cluster, but it is easily reproducible on *vanilla* Ceph with vstart as well, using the following steps:
$ git clone https://github.com/ceph/ceph.git && cd ceph $ ./do_cmake.sh -DCMAKE_BUILD_TYPE=RelWithDebInfo -DWITH_MANPAGE=OFF -DWITH_BABELTRACE=OFF -DWITH_MGR_DASHBOARD_FRONTEND=OFF && cd build && make -j32 vstart $ MON=1 OSD=1 MDS=0 ../src/vstart.sh --debug --new --localhost --bluestore --bluestore-devs /dev/xxx $ sudo ./bin/ceph osd pool create foo 32 32 $ sudo ./bin/rados bench -p foo 100 write --no-cleanup
With the old hard drive that I have (Hitachi HUA72201), I'm getting an average throughput of 60 MiB/s. When I switch to v15.2.4 (git checkout v15.2.4), rebuild, and repeat the experiment, and I get an average throughput of 90 MiB/s. I've reliably reproduced similar difference between 15.2.4 and HEAD by building release packages and running them on an 8-node cluster.
Is this expected or is this a performance regression?
Thanks!
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Hi Igor, If it is easily reproducible, could you run each test with gdbpmp so we can observe where time is being spent? Mark On 8/24/20 9:26 AM, Igor Fedotov wrote:
Hi All,
I've just completed a small investigation on this issue.
Short summary for Abutalib's benchmark scenario in my environment:
- when "vstarted" cluster has WAL/DB at NVMe drive - master is ~65% faster (130 MB/s vs. 75MB/s)
- in spinner-only setup (WAL/DB are at the same spinner) - octopus is 70 % faster than master ( 88MB/s vs. 52 MB/s)
Looks like it is bluestore_max_blob_size_hdd set to 64K by default which causes performance drop in this case (commit sha a8733598eddf57dca86bf002653e630a7cf4db6e).
Setting it back to 512K results in a better performance for master: 113 MB/s vs. 88 MB/s
But I presume these numbers are tightly bound to writing 4MB chunks benchmark scenario. Different chunk size will most probably cause pretty different results. Hence reverting bluestore_max_blob_size_hdd back to 512K by default is IMO still questionable.
Wondering if the above results are in-line with somebody's else ones....
Thanks,
Igor
On 8/4/2020 6:29 AM, Abutalib Aghayev wrote:
Hi all,
There appears to be a performance regression going from 15.2.4 to HEAD. I first realized this when testing my patches to Ceph on an 8-node cluster, but it is easily reproducible on *vanilla* Ceph with vstart as well, using the following steps:
$ git clone https://github.com/ceph/ceph.git && cd ceph $ ./do_cmake.sh -DCMAKE_BUILD_TYPE=RelWithDebInfo -DWITH_MANPAGE=OFF -DWITH_BABELTRACE=OFF -DWITH_MGR_DASHBOARD_FRONTEND=OFF && cd build && make -j32 vstart $ MON=1 OSD=1 MDS=0 ../src/vstart.sh --debug --new --localhost --bluestore --bluestore-devs /dev/xxx $ sudo ./bin/ceph osd pool create foo 32 32 $ sudo ./bin/rados bench -p foo 100 write --no-cleanup
With the old hard drive that I have (Hitachi HUA72201), I'm getting an average throughput of 60 MiB/s. When I switch to v15.2.4 (git checkout v15.2.4), rebuild, and repeat the experiment, and I get an average throughput of 90 MiB/s. I've reliably reproduced similar difference between 15.2.4 and HEAD by building release packages and running them on an 8-node cluster.
Is this expected or is this a performance regression?
Thanks!
_______________________________________________ Dev mailing list --dev@ceph.io To unsubscribe send an email todev-leave@ceph.io
Hrm, on second thought I suspect it's just more seeks on disk. Maybe blktrace would be more useful. Mark On 8/24/20 2:03 PM, Mark Nelson wrote:
Hi Igor,
If it is easily reproducible, could you run each test with gdbpmp so we can observe where time is being spent?
Mark
On 8/24/20 9:26 AM, Igor Fedotov wrote:
Hi All,
I've just completed a small investigation on this issue.
Short summary for Abutalib's benchmark scenario in my environment:
- when "vstarted" cluster has WAL/DB at NVMe drive - master is ~65% faster (130 MB/s vs. 75MB/s)
- in spinner-only setup (WAL/DB are at the same spinner) - octopus is 70 % faster than master ( 88MB/s vs. 52 MB/s)
Looks like it is bluestore_max_blob_size_hdd set to 64K by default which causes performance drop in this case (commit sha a8733598eddf57dca86bf002653e630a7cf4db6e).
Setting it back to 512K results in a better performance for master: 113 MB/s vs. 88 MB/s
But I presume these numbers are tightly bound to writing 4MB chunks benchmark scenario. Different chunk size will most probably cause pretty different results. Hence reverting bluestore_max_blob_size_hdd back to 512K by default is IMO still questionable.
Wondering if the above results are in-line with somebody's else ones....
Thanks,
Igor
On 8/4/2020 6:29 AM, Abutalib Aghayev wrote:
Hi all,
There appears to be a performance regression going from 15.2.4 to HEAD. I first realized this when testing my patches to Ceph on an 8-node cluster, but it is easily reproducible on *vanilla* Ceph with vstart as well, using the following steps:
$ git clone https://github.com/ceph/ceph.git && cd ceph $ ./do_cmake.sh -DCMAKE_BUILD_TYPE=RelWithDebInfo -DWITH_MANPAGE=OFF -DWITH_BABELTRACE=OFF -DWITH_MGR_DASHBOARD_FRONTEND=OFF && cd build && make -j32 vstart $ MON=1 OSD=1 MDS=0 ../src/vstart.sh --debug --new --localhost --bluestore --bluestore-devs /dev/xxx $ sudo ./bin/ceph osd pool create foo 32 32 $ sudo ./bin/rados bench -p foo 100 write --no-cleanup
With the old hard drive that I have (Hitachi HUA72201), I'm getting an average throughput of 60 MiB/s. When I switch to v15.2.4 (git checkout v15.2.4), rebuild, and repeat the experiment, and I get an average throughput of 90 MiB/s. I've reliably reproduced similar difference between 15.2.4 and HEAD by building release packages and running them on an 8-node cluster.
Is this expected or is this a performance regression?
Thanks!
_______________________________________________ Dev mailing list --dev@ceph.io To unsubscribe send an email todev-leave@ceph.io
yeah, that's my view as well... On 8/24/2020 10:07 PM, Mark Nelson wrote:
Hrm, on second thought I suspect it's just more seeks on disk. Maybe blktrace would be more useful.
Mark
On 8/24/20 2:03 PM, Mark Nelson wrote:
Hi Igor,
If it is easily reproducible, could you run each test with gdbpmp so we can observe where time is being spent?
Mark
On 8/24/20 9:26 AM, Igor Fedotov wrote:
Hi All,
I've just completed a small investigation on this issue.
Short summary for Abutalib's benchmark scenario in my environment:
- when "vstarted" cluster has WAL/DB at NVMe drive - master is ~65% faster (130 MB/s vs. 75MB/s)
- in spinner-only setup (WAL/DB are at the same spinner) - octopus is 70 % faster than master ( 88MB/s vs. 52 MB/s)
Looks like it is bluestore_max_blob_size_hdd set to 64K by default which causes performance drop in this case (commit sha a8733598eddf57dca86bf002653e630a7cf4db6e).
Setting it back to 512K results in a better performance for master: 113 MB/s vs. 88 MB/s
But I presume these numbers are tightly bound to writing 4MB chunks benchmark scenario. Different chunk size will most probably cause pretty different results. Hence reverting bluestore_max_blob_size_hdd back to 512K by default is IMO still questionable.
Wondering if the above results are in-line with somebody's else ones....
Thanks,
Igor
On 8/4/2020 6:29 AM, Abutalib Aghayev wrote:
Hi all,
There appears to be a performance regression going from 15.2.4 to HEAD. I first realized this when testing my patches to Ceph on an 8-node cluster, but it is easily reproducible on *vanilla* Ceph with vstart as well, using the following steps:
$ git clone https://github.com/ceph/ceph.git && cd ceph $ ./do_cmake.sh -DCMAKE_BUILD_TYPE=RelWithDebInfo -DWITH_MANPAGE=OFF -DWITH_BABELTRACE=OFF -DWITH_MGR_DASHBOARD_FRONTEND=OFF && cd build && make -j32 vstart $ MON=1 OSD=1 MDS=0 ../src/vstart.sh --debug --new --localhost --bluestore --bluestore-devs /dev/xxx $ sudo ./bin/ceph osd pool create foo 32 32 $ sudo ./bin/rados bench -p foo 100 write --no-cleanup
With the old hard drive that I have (Hitachi HUA72201), I'm getting an average throughput of 60 MiB/s. When I switch to v15.2.4 (git checkout v15.2.4), rebuild, and repeat the experiment, and I get an average throughput of 90 MiB/s. I've reliably reproduced similar difference between 15.2.4 and HEAD by building release packages and running them on an 8-node cluster.
Is this expected or is this a performance regression?
Thanks!
_______________________________________________ Dev mailing list --dev@ceph.io To unsubscribe send an email todev-leave@ceph.io
Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
To me it looks strange that max_blob_size isn't set to 4MB by default. Why isn't it? 24 августа 2020 г. 17:26:22 GMT+03:00, Igor Fedotov <ifedotov@suse.de> пишет:
Hi All,
I've just completed a small investigation on this issue.
Short summary for Abutalib's benchmark scenario in my environment:
- when "vstarted" cluster has WAL/DB at NVMe drive - master is ~65% faster (130 MB/s vs. 75MB/s)
- in spinner-only setup (WAL/DB are at the same spinner) - octopus is 70 % faster than master ( 88MB/s vs. 52 MB/s)
Looks like it is bluestore_max_blob_size_hdd set to 64K by default which causes performance drop in this case (commit sha a8733598eddf57dca86bf002653e630a7cf4db6e).
Setting it back to 512K results in a better performance for master: 113
MB/s vs. 88 MB/s
But I presume these numbers are tightly bound to writing 4MB chunks benchmark scenario. Different chunk size will most probably cause pretty different results. Hence reverting bluestore_max_blob_size_hdd back to 512K by default is IMO still questionable.
Wondering if the above results are in-line with somebody's else ones....
Thanks,
Igor
On 8/4/2020 6:29 AM, Abutalib Aghayev wrote:
Hi all,
There appears to be a performance regression going from 15.2.4 to HEAD. I first realized this when testing my patches to Ceph on an 8-node cluster, but it is easily reproducible on *vanilla* Ceph with vstart as well, using the following steps:
$ git clone https://github.com/ceph/ceph.git && cd ceph $ ./do_cmake.sh -DCMAKE_BUILD_TYPE=RelWithDebInfo -DWITH_MANPAGE=OFF -DWITH_BABELTRACE=OFF -DWITH_MGR_DASHBOARD_FRONTEND=OFF && cd build && make -j32 vstart $ MON=1 OSD=1 MDS=0 ../src/vstart.sh --debug --new --localhost --bluestore --bluestore-devs /dev/xxx $ sudo ./bin/ceph osd pool create foo 32 32 $ sudo ./bin/rados bench -p foo 100 write --no-cleanup
With the old hard drive that I have (Hitachi HUA72201), I'm getting an average throughput of 60 MiB/s. When I switch to v15.2.4 (git checkout v15.2.4), rebuild, and repeat the experiment, and I get an average throughput of 90 MiB/s. I've reliably reproduced similar difference between 15.2.4 and HEAD by building release packages and running them on an 8-node cluster.
Is this expected or is this a performance regression?
Thanks!
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
-- With best regards, Vitaliy Filippov
participants (4)
-
Abutalib Aghayev
-
Igor Fedotov
-
Mark Nelson
-
Виталий Филиппов