RBD snapshots very slow
Hi. For a long time I was under an impression that clones are as efficient in bluestore as snapshots. But today I finally decided to test it and ... I discovered it was an utterly wrong impression :) RBD copies the whole 4 MB object even when a small 4 KB block is modified within it in the child image. In my all-NVMe cluster this leads to 40 (40!!!) random write iops (bs=4k iodepth=1) in a fresh RBD clone, which is terrible. Question of the day: is it possible to reimplement RBD clones using "sparse objects"? As I understand the support for sparse objects themselves is already there. So maybe librbd could only write the modified part to the child image when writing and read "holes" from parents when reading? -- Vitaliy Filippov
On Fri, Mar 20, 2020 at 1:29 PM <vitalif@yourcmc.ru> wrote:
Hi.
For a long time I was under an impression that clones are as efficient in bluestore as snapshots.
But today I finally decided to test it and ... I discovered it was an utterly wrong impression :) RBD copies the whole 4 MB object even when a small 4 KB block is modified within it in the child image. In my all-NVMe cluster this leads to 40 (40!!!) random write iops (bs=4k iodepth=1) in a fresh RBD clone, which is terrible.
Anything with an iodepth of 1 is going to be (relatively) terrible on RBD.
Question of the day: is it possible to reimplement RBD clones using "sparse objects"? As I understand the support for sparse objects themselves is already there. So maybe librbd could only write the modified part to the child image when writing and read "holes" from parents when reading?
The forthcoming Octopus release of librbd adds support for sparse copy-up writes [1] when your min OSD release is set to Octopus (reads from the parent image were already sparse-read ops). Using holes was previously not very practical due to the large allocations sizes on the OSD, but with the change to 4KiB minimum block sizes, such a technique would be possible (albeit a breaking change for all older clients controlled via a new feature bit). You also have the ability to control the RBD object sizes and use something smaller than the 4MiB default.
-- Vitaliy Filippov _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
[1] https://github.com/ceph/ceph/pull/27999 -- Jason
Anything with an iodepth of 1 is going to be (relatively) terrible on RBD.
iodepth=128 results in 300 iops in the same setup. Not much better :)
The forthcoming Octopus release of librbd adds support for sparse copy-up writes [1] when your min OSD release is set to Octopus (reads from the parent image were already sparse-read ops). Using holes was previously not very practical due to the large allocations sizes on the OSD, but with the change to 4KiB minimum block sizes, such a technique would be possible (albeit a breaking change for all older clients controlled via a new feature bit). You also have the ability to control the RBD object sizes and use something smaller than the 4MiB default.
That's good news! Is it already in 15.1.1, can I test it? -- Vitaliy Filippov
On Fri, Mar 20, 2020 at 2:23 PM <vitalif@yourcmc.ru> wrote:
Anything with an iodepth of 1 is going to be (relatively) terrible on RBD.
iodepth=128 results in 300 iops in the same setup. Not much better :)
The forthcoming Octopus release of librbd adds support for sparse copy-up writes [1] when your min OSD release is set to Octopus (reads from the parent image were already sparse-read ops). Using holes was previously not very practical due to the large allocations sizes on the OSD, but with the change to 4KiB minimum block sizes, such a technique would be possible (albeit a breaking change for all older clients controlled via a new feature bit). You also have the ability to control the RBD object sizes and use something smaller than the 4MiB default.
That's good news! Is it already in 15.1.1, can I test it?
Yup, but you need 15.1.z OSDs as well as 15.1.z librbd. Also ensure you run "ceph osd require-osd-release octopus" otherwise it won't use the new sparse-copyup since older OSDs don't have the necessary support.
-- Vitaliy Filippov
-- Jason
Yup, but you need 15.1.z OSDs as well as 15.1.z librbd. Also ensure you run "ceph osd require-osd-release octopus" otherwise it won't use the new sparse-copyup since older OSDs don't have the necessary support.
OK, I've just upgraded my test cluster to 15.1.1 and tried to test this feature. It didn't seem to work. I tested it with: fio -ioengine=rbd -direct=1 -name=test -bs=64k -iodepth=1 -rw=randwrite -pool=rpool_hdd -rbdname=testimg3_clone The result was 10 iops while `ceph -s` reported 40 MB/s write... Just in case if it depends on min_alloc_size I've also tested it with bs=64k with the same result. Is there anything else I should do to get it working? -- Vitaliy Filippov
Oh, sorry, you probably didn't mean that it's the speedup that was already implemented, just the sparse writes into child images... OK... -- Vitaliy Filippov
On Fri, Mar 20, 2020 at 01:41:11PM -0400, Jason Dillaman wrote:
The forthcoming Octopus release of librbd adds support for sparse copy-up writes [1] when your min OSD release is set to Octopus (reads from the parent image were already sparse-read ops). Using holes was previously not very practical due to the large allocations sizes on the OSD, but with the change to 4KiB minimum block sizes, such a technique would be possible (albeit a breaking change for all older clients controlled via a new feature bit). You also have the ability to control the RBD object sizes and use something smaller than the 4MiB default.
Also note, that EC pools do not support sparse reads, which makes it a bit less useful, taking into account how common EC pools become now. -- Mykola Golub
participants (3)
-
Jason Dillaman
-
Mykola Golub
-
vitalif@yourcmc.ru