Hi, We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold. The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286 tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed. Any advice and suggestions will be greatly appreciated. thx, Yang Honggang
Hi Honggang, I personally I find this very exciting! I was hoping that we might eventually try local caching in bluestore especially given trends for larger NVMe devices and pmem. When you were running performance tests, did you run any tests where the data set size was significantly larger than the available "fast" local tier cache (ie so that eviction was taking place)? In the past, that's been the area we've really needed to focus on getting right. Mark On 10/11/19 11:04 AM, Honggang(Joseph) Yang wrote:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang
After the sysbench prepare operation is completed, about 48883MB of db data is generated. I set the fast partition to 30GB, so in the sysbench run stage, eviction was taking place. Mark Nelson <mnelson@redhat.com> 于2019年10月12日周六 上午12:15写道:
Hi Honggang,
I personally I find this very exciting! I was hoping that we might eventually try local caching in bluestore especially given trends for larger NVMe devices and pmem. When you were running performance tests, did you run any tests where the data set size was significantly larger than the available "fast" local tier cache (ie so that eviction was taking place)? In the past, that's been the area we've really needed to focus on getting right.
Mark
On 10/11/19 11:04 AM, Honggang(Joseph) Yang wrote:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang
Several considerations on deployment: 1. How a traditional OSD(filestore, or bluestore W/O local cache) behavior if the OSD is selected as a PG member of the local-cache-pool? hopefully the osd can just work as-is like an OSD with CACHE-FULL. 2. Compared to other *global* caching mode, local-cache-mode cannot scale cache tier/storage tier separately, which will be a big caveats in real production environment, especially for those case that the working set is consistent but the archival data amount is keeping increasing. Overall it seems like a alternative for b-cache/flashcache/i-cas , more test or analysis is needed to understand the benefit of this approach over the existing b-cache solution. Honggang(Joseph) Yang <eagle.rtlinux@gmail.com> 于2019年10月13日周日 下午8:47写道:
After the sysbench prepare operation is completed, about 48883MB of db data is generated. I set the fast partition to 30GB, so in the sysbench run stage, eviction was taking place.
Mark Nelson <mnelson@redhat.com> 于2019年10月12日周六 上午12:15写道:
Hi Honggang,
I personally I find this very exciting! I was hoping that we might eventually try local caching in bluestore especially given trends for larger NVMe devices and pmem. When you were running performance tests, did you run any tests where the data set size was significantly larger than the available "fast" local tier cache (ie so that eviction was taking place)? In the past, that's been the area we've really needed to focus on getting right.
Mark
On 10/11/19 11:04 AM, Honggang(Joseph) Yang wrote:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode:
https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang
On Mon, 14 Oct 2019 at 10:31, Xiaoxi Chen <superdebuger@gmail.com> wrote:
Several considerations on deployment:
1. How a traditional OSD(filestore, or bluestore W/O local cache) behavior if the OSD is selected as a PG member of the local-cache-pool? hopefully the osd can just work as-is like an OSD with CACHE-FULL.
Only the pools whose cache mode is local will be affected, all other pools share the same osds will just act like before.
2. Compared to other *global* caching mode, local-cache-mode cannot scale cache tier/storage tier separately, which will be a big caveats in real production environment, especially for those case that the working set is consistent but the archival data amount is keeping increasing.
Overall it seems like a alternative for b-cache/flashcache/i-cas , more test or analysis is needed to understand the benefit of this approach over the existing b-cache solution.
Yes, this is a problem. For now, tier local mode only support SSD:HDD = 1:1 mode. If you want to extend cache's size under tier local mode, you can: - disable tier local mode - flush all fast objects to HDD - replace SSD dev to a large one - enable local mode again or add more SSD:HDD pairs into the related crush domain.
Honggang(Joseph) Yang <eagle.rtlinux@gmail.com> 于2019年10月13日周日 下午8:47写道:
After the sysbench prepare operation is completed, about 48883MB of db data is generated. I set the fast partition to 30GB, so in the sysbench run stage, eviction was taking place.
Mark Nelson <mnelson@redhat.com> 于2019年10月12日周六 上午12:15写道:
Hi Honggang,
I personally I find this very exciting! I was hoping that we might eventually try local caching in bluestore especially given trends for larger NVMe devices and pmem. When you were running performance tests, did you run any tests where the data set size was significantly larger than the available "fast" local tier cache (ie so that eviction was taking place)? In the past, that's been the area we've really needed to focus on getting right.
Mark
On 10/11/19 11:04 AM, Honggang(Joseph) Yang wrote:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang
Looks quite interesting, i do however think local caching is better done at the block level (bcache, dm-cache, dm-writecache) rather than in Ceph. In theory they can deal with a smaller granularity than a Ceph object + go through the kernel block layer which is more optimized than a Ceph OSD. Your results do show favorable comparison with bcache, it will be good to try to know why this is the case (at least at a high level), i know cache testing/simulation is not easy to compare two caching methods, but i think it is important to know why local Ceph caching would be better. It will also be interesting to compare it with dm-writecache, which is optimized for writes (using pmem or ssd devices) which is in many cases the main performance bottleneck as reads can be cached in memory (assuming you have enough ram). So i think more tests need to be done, which for caching is not a simple matter. I believe fio does have a random_distribution=zipf:[theta] parameter trying to simulate semi real io, as pure serial or pure random io is not suitable for testing cache. /Maged On 11/10/2019 18:04, Honggang(Joseph) Yang wrote:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Sat, 12 Oct 2019 at 04:58, Maged Mokhtar <mmokhtar@petasan.org> wrote:
Looks quite interesting, i do however think local caching is better done at the block level (bcache, dm-cache, dm-writecache) rather than in Ceph. In theory they can deal with a smaller granularity than a Ceph object + go through the kernel block layer which is more optimized than a Ceph OSD.
yes, fine granularity and low migration cost. But, based on local mode tier, we can implement more flexible migration strategies. Take cephfs as an example, there are a lot of files stored in the system, but we only need to edit some of them in a certain period of time. We can migrate the data part of the related files to SSD before the editing process through the hint operation. This ensures that the most important data is in SSD, while other files still stay on HDD. It is not easy to do this kind of work based on block level cache implementation, because they don't know the up layer logical objects.
Your results do show favorable comparison with bcache, it will be good to try to know why this is the case (at least at a high level), i know cache testing/simulation is not easy to compare two caching methods, but i think it is important to know why local Ceph caching would be better.
It will also be interesting to compare it with dm-writecache, which is optimized for writes (using pmem or ssd devices) which is in many cases the main performance bottleneck as reads can be cached in memory (assuming you have enough ram).
So i think more tests need to be done, which for caching is not a simple matter. I believe fio does have a random_distribution=zipf:[theta]
I made a comparison with CAS: https://tracker.ceph.com/issues/42286?next_issue_id=42285#I-also-compared-lo...
parameter trying to simulate semi real io, as pure serial or pure random io is not suitable for testing cache.
As for local's performance, it's highly related to how to figure out hot object. Now I reuse pool tier's hitset to do this work. If the migration overhead can be compensated by subsequent read and write hits, performance is not a problem.
/Maged
On 11/10/2019 18:04, Honggang(Joseph) Yang wrote:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Honggang(Joseph) Yang <eagle.rtlinux@gmail.com> 于2019年10月12日周六 上午12:04写道:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
I'm quite interesting about this and seems quite promising. But it seems objects level tier seems still not efficient enough in the small write scenario (4k ~ 8k based on your fio test result), right? Now, in bluestore, it is possbile to implement a fine-grain local tier. it seems bluestore_pextent_t may be allocated from different devices and strategies like extent level hitset to evaluate whether some extents need to demote to slow devices. A promotion or demotion may submit a read with write operations in a transactions, and finished with a onode or blobs flushed into rocksdb. The introduction of tier local mode in detail is
https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
This is quite interesting. I'll have a closer look this week and get back to you with questions! -Sam On Tue, Oct 15, 2019 at 5:15 AM Ning Yao <zay11022@gmail.com> wrote:
Honggang(Joseph) Yang <eagle.rtlinux@gmail.com> 于2019年10月12日周六 上午12:04写道:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
I'm quite interesting about this and seems quite promising. But it seems objects level tier seems still not efficient enough in the small write scenario (4k ~ 8k based on your fio test result), right? Now, in bluestore, it is possbile to implement a fine-grain local tier. it seems bluestore_pextent_t may be allocated from different devices and strategies like extent level hitset to evaluate whether some extents need to demote to slow devices. A promotion or demotion may submit a read with write operations in a transactions, and finished with a onode or blobs flushed into rocksdb.
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Thank you. I look forward to hearing from you. On Wed, 16 Oct 2019 at 04:24, Sam Just <sjust@redhat.com> wrote:
This is quite interesting. I'll have a closer look this week and get back to you with questions! -Sam
On Tue, Oct 15, 2019 at 5:15 AM Ning Yao <zay11022@gmail.com> wrote:
Honggang(Joseph) Yang <eagle.rtlinux@gmail.com> 于2019年10月12日周六 上午12:04写道:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
I'm quite interesting about this and seems quite promising. But it seems objects level tier seems still not efficient enough in the small write scenario (4k ~ 8k based on your fio test result), right? Now, in bluestore, it is possbile to implement a fine-grain local tier. it seems bluestore_pextent_t may be allocated from different devices and strategies like extent level hitset to evaluate whether some extents need to demote to slow devices. A promotion or demotion may submit a read with write operations in a transactions, and finished with a onode or blobs flushed into rocksdb.
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
_______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On Tue, 15 Oct 2019 at 20:15, Ning Yao <zay11022@gmail.com> wrote:
Honggang(Joseph) Yang <eagle.rtlinux@gmail.com> 于2019年10月12日周六 上午12:04写道:
Hi,
We implemented a new cache tier mode - local mode. In this mode, an osd is configured to manage two data devices, one is fast device, one is slow device. Hot objects are promoted from slow device to fast device, and demoted from fast device to slow device when they become cold.
I'm quite interesting about this and seems quite promising. But it seems objects level tier seems still not efficient enough in the small write scenario (4k ~ 8k based on your fio test result), right? Now, in bluestore, it is possbile to implement a fine-grain local tier. it seems bluestore_pextent_t may be allocated from different devices and strategies like extent level hitset to evaluate whether some extents need to demote to slow devices. A promotion or demotion may submit a read with write operations in a transactions, and finished with a onode or blobs flushed into rocksdb.
Tier local mode offers a more flexible mechanism for the up layer(rgw/cephfs/rbd) to participate the tier migration work. As for the performance test result, its highly dependent on the io pattern and quality of my code :)
The introduction of tier local mode in detail is https://tracker.ceph.com/issues/42286
tier local mode: https://github.com/yanghonggang/ceph/commits/wip-tier-new
This work is based on ceph v12.2.5. I'm glad to port it to master branch if needed.
Any advice and suggestions will be greatly appreciated.
thx,
Yang Honggang _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
On 2019-10-16T11:28:25, " Honggang(Joseph) Yang " <eagle.rtlinux@gmail.com> wrote: I'm glad to see more performance work and caching happening in Ceph! I admit calling this a "tier" (to get the bike shedding done first ;-) is confusing me, because that used to mean something different. This seems to me to be more of a BlueStore feature based on hints/access from the upper layers? So perhaps, at that level, it'd make sense to instead use the space on the RocksDB partition/device for this caching operation, instead of yet an additional device? (Intuitively, that's what most users already expect it does, anyway.) How would this, compared to bcache, possibly handle situations where multiple OSDs share one caching device? And does this only promote the local shard/replica? I'm wondering how this would affect EC pools. Regards, Lars -- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
On Thu, 17 Oct 2019 at 06:22, Lars Marowsky-Bree <lmb@suse.com> wrote:
On 2019-10-16T11:28:25, " Honggang(Joseph) Yang " <eagle.rtlinux@gmail.com> wrote:
I'm glad to see more performance work and caching happening in Ceph!
I admit calling this a "tier" (to get the bike shedding done first ;-) is confusing me, because that used to mean something different. This seems to me to be more of a BlueStore feature based on hints/access from the upper layers?
User can explicitly send hint op or the do_op/agent send hint op based on object access statistics to trigger a migration.
So perhaps, at that level, it'd make sense to instead use the space on the RocksDB partition/device for this caching operation, instead of yet an additional device? (Intuitively, that's what most users already expect it does, anyway.)
yes, this is user friendly.
How would this, compared to bcache, possibly handle situations where multiple OSDs share one caching device?
SSD is split into multiple partitions. Each partition is assigned to an osd as fast partitions.
And does this only promote the local shard/replica? I'm wondering how this would affect EC pools.
yes, only promote the local shard/replica. But there is still some work to do to support ec pool.
Regards, Lars
-- SUSE Linux GmbH, GF: Felix Imendörffer, Mary Higgins, Sri Rasiah, HRB 21284 (AG Nürnberg) "Architects should open possibilities and not determine everything." (Ueli Zbinden)
participants (7)
-
Honggang(Joseph) Yang
-
Lars Marowsky-Bree
-
Maged Mokhtar
-
Mark Nelson
-
Ning Yao
-
Sam Just
-
Xiaoxi Chen