block.db/block.wal device performance dropped after upgrade to 14.2.10
Good day, cephers! We've recently upgraded our cluster from 14.2.8 to 14.2.10 release, also performing full system packages upgrade(Ubuntu 18.04 LTS). After that performance significantly dropped, main reason beeing that journal SSDs are now have no merges, huge queues, and increased latency. There's a few screenshots in attachments. This is for an SSD journal that supports block.db/block.wal for 3 spinning OSDs, and it looks like this for all our SSD block.db/wal devices across all nodes. Any ideas what may cause that? Maybe I've missed something important in release notes?
Here's some more insight into the issue. Looks like the load is triggered because of a snaptrim operation. We have a backup pool that serves as Openstack cinder-backup storage, performing snapshot backups every night. Old backups are also deleted every night, so snaptrim is initiated. This snaptrim increased load on the block.db devices after upgrade, and just kills one SSD's performance in particular. It serves as a block.db/wal device for one of the fatter backup pool OSDs which has more PGs placed there. This is a Kingston SSD, and we see this issue on other Kingston SSD journals too, Intel SSD journals are not that affected, though they too experience increased load. Nevertheless, there're now a lot of read IOPS on block.db devices after upgrade that were not there before. I wonder how 600 IOPS can destroy SSDs performance that hard. вт, 4 авг. 2020 г. в 12:54, Vladimir Prokofev <v@prokofev.me>:
Good day, cephers!
We've recently upgraded our cluster from 14.2.8 to 14.2.10 release, also performing full system packages upgrade(Ubuntu 18.04 LTS). After that performance significantly dropped, main reason beeing that journal SSDs are now have no merges, huge queues, and increased latency. There's a few screenshots in attachments. This is for an SSD journal that supports block.db/block.wal for 3 spinning OSDs, and it looks like this for all our SSD block.db/wal devices across all nodes. Any ideas what may cause that? Maybe I've missed something important in release notes?
Hi Vladimir, What Kingston SSD model? El 4/8/20 a las 12:22, Vladimir Prokofev escribió:
Here's some more insight into the issue. Looks like the load is triggered because of a snaptrim operation. We have a backup pool that serves as Openstack cinder-backup storage, performing snapshot backups every night. Old backups are also deleted every night, so snaptrim is initiated. This snaptrim increased load on the block.db devices after upgrade, and just kills one SSD's performance in particular. It serves as a block.db/wal device for one of the fatter backup pool OSDs which has more PGs placed there. This is a Kingston SSD, and we see this issue on other Kingston SSD journals too, Intel SSD journals are not that affected, though they too experience increased load. Nevertheless, there're now a lot of read IOPS on block.db devices after upgrade that were not there before. I wonder how 600 IOPS can destroy SSDs performance that hard.
вт, 4 авг. 2020 г. в 12:54, Vladimir Prokofev <v@prokofev.me>:
Good day, cephers!
We've recently upgraded our cluster from 14.2.8 to 14.2.10 release, also performing full system packages upgrade(Ubuntu 18.04 LTS). After that performance significantly dropped, main reason beeing that journal SSDs are now have no merges, huge queues, and increased latency. There's a few screenshots in attachments. This is for an SSD journal that supports block.db/block.wal for 3 spinning OSDs, and it looks like this for all our SSD block.db/wal devices across all nodes. Any ideas what may cause that? Maybe I've missed something important in release notes?
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Eneko Lacunza | Tel. 943 569 206 | Email elacunza@binovo.es Director Técnico | Site. https://www.binovo.es BINOVO IT HUMAN PROJECT S.L | Dir. Astigarragako Bidea, 2 - 2º izda. Oficina 10-11, 20180 Oiartzun
What Kingston SSD model?
=== START OF INFORMATION SECTION === Model Family: SandForce Driven SSDs Device Model: KINGSTON SE50S3100G Serial Number: xxxxxxxxxxxxxxxx LU WWN Device Id: xxxxxxxxxxxxxxxx Firmware Version: 611ABBF0 User Capacity: 100,030,242,816 bytes [100 GB] Sector Size: 512 bytes logical/physical Rotation Rate: Solid State Device Form Factor: 2.5 inches Device is: In smartctl database [for details use: -P show] ATA Version is: ATA8-ACS, ACS-2 T13/2015-D revision 3 SATA Version is: SATA 3.0, 6.0 Gb/s (current: 6.0 Gb/s) Local Time is: Tue Aug 4 14:31:36 2020 MSK SMART support is: Available - device has SMART capability. SMART support is: Enabled вт, 4 авг. 2020 г. в 14:17, Eneko Lacunza <elacunza@binovo.es>:
Hi Vladimir,
What Kingston SSD model?
Here's some more insight into the issue. Looks like the load is triggered because of a snaptrim operation. We have a backup pool that serves as Openstack cinder-backup storage, performing snapshot backups every night. Old backups are also deleted every night, so snaptrim is initiated. This snaptrim increased load on the block.db devices after upgrade, and just kills one SSD's performance in particular. It serves as a block.db/wal device for one of the fatter backup pool OSDs which has more PGs placed there. This is a Kingston SSD, and we see this issue on other Kingston SSD journals too, Intel SSD journals are not that affected, though they too experience increased load. Nevertheless, there're now a lot of read IOPS on block.db devices after upgrade that were not there before. I wonder how 600 IOPS can destroy SSDs performance that hard.
вт, 4 авг. 2020 г. в 12:54, Vladimir Prokofev <v@prokofev.me>:
Good day, cephers!
We've recently upgraded our cluster from 14.2.8 to 14.2.10 release, also performing full system packages upgrade(Ubuntu 18.04 LTS). After that performance significantly dropped, main reason beeing that journal SSDs are now have no merges, huge queues, and increased latency. There's a few screenshots in attachments. This is for an SSD journal
El 4/8/20 a las 12:22, Vladimir Prokofev escribió: that
supports block.db/block.wal for 3 spinning OSDs, and it looks like this for all our SSD block.db/wal devices across all nodes. Any ideas what may cause that? Maybe I've missed something important in release notes?
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Eneko Lacunza | Tel. 943 569 206 | Email elacunza@binovo.es Director Técnico | Site. https://www.binovo.es BINOVO IT HUMAN PROJECT S.L | Dir. Astigarragako Bidea, 2 - 2º izda. Oficina 10-11, 20180 Oiartzun _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I really would not focus that much on a particular device model. Yes, Kingston SSDs are slower for reads, we knew that since we tested them. But that was before they were used as block.db devices, they first were intended purely as block.wal devices. This was even before bluestore actually, so their primary function was to serve as XFS journals. I see increased read activity on all block.db devices while snaptrims are active, be them Intel or Kingston. Latter just suffer more. This read load was not there before 14.2.10/system upgrade. I'm actually thinking about rolling back to 14.2.8. Any ideas how safe that procedure is? I suppose it should be safe since there was no change in the actual data storage scheme? вт, 4 авг. 2020 г. в 14:33, Vladimir Prokofev <v@prokofev.me>:
What Kingston SSD model?
=== START OF INFORMATION SECTION === Model Family: SandForce Driven SSDs Device Model: KINGSTON SE50S3100G Serial Number: xxxxxxxxxxxxxxxx LU WWN Device Id: xxxxxxxxxxxxxxxx Firmware Version: 611ABBF0 User Capacity: 100,030,242,816 bytes [100 GB] Sector Size: 512 bytes logical/physical Rotation Rate: Solid State Device Form Factor: 2.5 inches Device is: In smartctl database [for details use: -P show] ATA Version is: ATA8-ACS, ACS-2 T13/2015-D revision 3 SATA Version is: SATA 3.0, 6.0 Gb/s (current: 6.0 Gb/s) Local Time is: Tue Aug 4 14:31:36 2020 MSK SMART support is: Available - device has SMART capability. SMART support is: Enabled
вт, 4 авг. 2020 г. в 14:17, Eneko Lacunza <elacunza@binovo.es>:
Hi Vladimir,
What Kingston SSD model?
Here's some more insight into the issue. Looks like the load is triggered because of a snaptrim operation. We have a backup pool that serves as Openstack cinder-backup storage, performing snapshot backups every night. Old backups are also deleted every night, so snaptrim is initiated. This snaptrim increased load on the block.db devices after upgrade, and just kills one SSD's performance in particular. It serves as a block.db/wal device for one of the fatter backup pool OSDs which has more PGs placed there. This is a Kingston SSD, and we see this issue on other Kingston SSD journals too, Intel SSD journals are not that affected, though they too experience increased load. Nevertheless, there're now a lot of read IOPS on block.db devices after upgrade that were not there before. I wonder how 600 IOPS can destroy SSDs performance that hard.
вт, 4 авг. 2020 г. в 12:54, Vladimir Prokofev <v@prokofev.me>:
Good day, cephers!
We've recently upgraded our cluster from 14.2.8 to 14.2.10 release, also performing full system packages upgrade(Ubuntu 18.04 LTS). After that performance significantly dropped, main reason beeing that journal SSDs are now have no merges, huge queues, and increased latency. There's a few screenshots in attachments. This is for an SSD journal
supports block.db/block.wal for 3 spinning OSDs, and it looks like
El 4/8/20 a las 12:22, Vladimir Prokofev escribió: that this for
all our SSD block.db/wal devices across all nodes. Any ideas what may cause that? Maybe I've missed something important in release notes?
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Eneko Lacunza | Tel. 943 569 206 | Email elacunza@binovo.es Director Técnico | Site. https://www.binovo.es BINOVO IT HUMAN PROJECT S.L | Dir. Astigarragako Bidea, 2 - 2º izda. Oficina 10-11, 20180 Oiartzun _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hello Vladimir, I just tested this with a single node testcluster with 60 HDDs (3 of them with bluestore without separate wal and db). With the 14.2.10, I see on the bluestore OSDs a lot of read IOPs while snaptrimming. With 14.2.9 this was not an issue. I wonder if this would explain the huge amount of slowops on my big testcluster (44 Nodes 1056 OSDs) while snaptrimming. I cannot test a downgrade there, because there are no packages of older releases for CentOS 8 available. Regards Manuel On Tue, 4 Aug 2020 13:22:34 +0300 Vladimir Prokofev <v@prokofev.me> wrote:
Here's some more insight into the issue. Looks like the load is triggered because of a snaptrim operation. We have a backup pool that serves as Openstack cinder-backup storage, performing snapshot backups every night. Old backups are also deleted every night, so snaptrim is initiated. This snaptrim increased load on the block.db devices after upgrade, and just kills one SSD's performance in particular. It serves as a block.db/wal device for one of the fatter backup pool OSDs which has more PGs placed there. This is a Kingston SSD, and we see this issue on other Kingston SSD journals too, Intel SSD journals are not that affected, though they too experience increased load. Nevertheless, there're now a lot of read IOPS on block.db devices after upgrade that were not there before. I wonder how 600 IOPS can destroy SSDs performance that hard.
вт, 4 авг. 2020 г. в 12:54, Vladimir Prokofev <v@prokofev.me>:
Good day, cephers!
We've recently upgraded our cluster from 14.2.8 to 14.2.10 release, also performing full system packages upgrade(Ubuntu 18.04 LTS). After that performance significantly dropped, main reason beeing that journal SSDs are now have no merges, huge queues, and increased latency. There's a few screenshots in attachments. This is for an SSD journal that supports block.db/block.wal for 3 spinning OSDs, and it looks like this for all our SSD block.db/wal devices across all nodes. Any ideas what may cause that? Maybe I've missed something important in release notes?
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi, I found the reasen of this behavior change. With 14.2.10 the default value of "bluefs_buffered_io" was changed from true to false. https://tracker.ceph.com/issues/44818 configureing this to true my problems seems to be solved. Regards Manuel On Wed, 5 Aug 2020 13:30:45 +0200 Manuel Lausch <manuel.lausch@1und1.de> wrote:
Hello Vladimir,
I just tested this with a single node testcluster with 60 HDDs (3 of them with bluestore without separate wal and db).
With the 14.2.10, I see on the bluestore OSDs a lot of read IOPs while snaptrimming. With 14.2.9 this was not an issue.
I wonder if this would explain the huge amount of slowops on my big testcluster (44 Nodes 1056 OSDs) while snaptrimming. I cannot test a downgrade there, because there are no packages of older releases for CentOS 8 available.
Regards Manuel
Maneul, thank you for your input. This is actually huge, and the problem is exactly that. On a side note I will add, that I observed lower memory utilisation on OSD nodes since the update, and a big throughput on block.db devices(up to 100+MB/s) that was not there before, so logically that meant that some operations that were performed in memory before, now were executed directly on block device. Was digging through possible causes, but your time-saving message arrived earlier. Thank you! чт, 6 авг. 2020 г. в 14:56, Manuel Lausch <manuel.lausch@1und1.de>:
Hi,
I found the reasen of this behavior change. With 14.2.10 the default value of "bluefs_buffered_io" was changed from true to false. https://tracker.ceph.com/issues/44818
configureing this to true my problems seems to be solved.
Regards Manuel
On Wed, 5 Aug 2020 13:30:45 +0200 Manuel Lausch <manuel.lausch@1und1.de> wrote:
Hello Vladimir,
I just tested this with a single node testcluster with 60 HDDs (3 of them with bluestore without separate wal and db).
With the 14.2.10, I see on the bluestore OSDs a lot of read IOPs while snaptrimming. With 14.2.9 this was not an issue.
I wonder if this would explain the huge amount of slowops on my big testcluster (44 Nodes 1056 OSDs) while snaptrimming. I cannot test a downgrade there, because there are no packages of older releases for CentOS 8 available.
Regards Manuel
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Thinking about this a little more, one thing that I remember when I was writing the priority cache manager is that in some cases I saw strange behavior with the rocksdb block cache when compaction was performed. It appeared that the entire contents of the cache could be invalidated. I guess that would only make sense if it was trimming old entries from the cache instead of entries associated with (now deleted) sst files or perhaps waiting to delete all SST files until the end of the compaction cycle thus forcing old entries out of the cache and then invalidating the whole works. In any event, I wonder if having the secondary page cache is enough on your clusters to sort of get around all of this by still having the SST files associated with the previously heavily used blocks in page cache kicking around until compaction completes. Maybe the combination of snap trimming or other background work along with compaction is just totally thrashing the rocksdb block cache. For folks that feel comfortable watching IO hitting your DB devices, can you see if you have increased bursts of reads to the DB device after a compaction event has occurred? They look like this in the OSD logs: 2020-08-04T17:15:56.603+0000 7fb0cf60d700 4 rocksdb: (Original Log Time 2020/08/04-17:15:56.603585) EVENT_LOG_v1 {"time_micros": 1596561356603574, "job": 5, "event": "compaction_finished", "compaction_time_micros": 744532, "compaction_time_cpu_micros": 607655, "output_level": 1, "num_output_files": 2, "total_output_size": 84712923, "num_input_records": 1714260, "num_output_records": 658541, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 2, 0, 0, 0, 0, 0]} You can also run this tool to get a nicely formatted list of them, though I don't have it reporting timestamps, just the time offset from the start of the log so looking at the OSD logs directly would be easier to match up timestamps. https://github.com/ceph/cbt/blob/master/tools/ceph_rocksdb_log_parser.py Mark On 8/6/20 8:07 AM, Vladimir Prokofev wrote:
Maneul, thank you for your input. This is actually huge, and the problem is exactly that.
On a side note I will add, that I observed lower memory utilisation on OSD nodes since the update, and a big throughput on block.db devices(up to 100+MB/s) that was not there before, so logically that meant that some operations that were performed in memory before, now were executed directly on block device. Was digging through possible causes, but your time-saving message arrived earlier. Thank you!
чт, 6 авг. 2020 г. в 14:56, Manuel Lausch <manuel.lausch@1und1.de>:
Hi,
I found the reasen of this behavior change. With 14.2.10 the default value of "bluefs_buffered_io" was changed from true to false. https://tracker.ceph.com/issues/44818
configureing this to true my problems seems to be solved.
Regards Manuel
On Wed, 5 Aug 2020 13:30:45 +0200 Manuel Lausch <manuel.lausch@1und1.de> wrote:
Hello Vladimir,
I just tested this with a single node testcluster with 60 HDDs (3 of them with bluestore without separate wal and db).
With the 14.2.10, I see on the bluestore OSDs a lot of read IOPs while snaptrimming. With 14.2.9 this was not an issue.
I wonder if this would explain the huge amount of slowops on my big testcluster (44 Nodes 1056 OSDs) while snaptrimming. I cannot test a downgrade there, because there are no packages of older releases for CentOS 8 available.
Regards Manuel
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Yeah, there are cases where enabling it will improve performance as rocksdb can then used the page cache as a (potentially large) secondary cache beyond the block cache and avoid hitting the underlying devices for reads. Do you have a lot of spare memory for page cache on your OSD nodes? You may be able to improve the situation with bluefs_buffered_io=false by increasing the osd_memory_target which should give the rocksdb block cache more memory to work with directly. One downside is that we currently double cache onodes in both the rocksdb cache and bluestore onode cache which hurts us when memory limited. We have some experimental work that might help in this area by better balancing bluestore onode and rocksdb block caches but it needs to be rebased after Adam's column family sharding work. The reason we had to disable bluefs_buffered_io again was that we had users with certain RGW workloads where the kernel started swapping large amounts of memory on the OSD nodes despite seemingly have free memory available. This caused huge latency spikes and IO slowdowns (even stalls). We never noticed it in our QA test suites and it doesn't appear to happen with RBD workloads as far as I can tell, but when it does happen it's really painful. Mark On 8/6/20 6:53 AM, Manuel Lausch wrote:
Hi,
I found the reasen of this behavior change. With 14.2.10 the default value of "bluefs_buffered_io" was changed from true to false. https://tracker.ceph.com/issues/44818
configureing this to true my problems seems to be solved.
Regards Manuel
On Wed, 5 Aug 2020 13:30:45 +0200 Manuel Lausch <manuel.lausch@1und1.de> wrote:
Hello Vladimir,
I just tested this with a single node testcluster with 60 HDDs (3 of them with bluestore without separate wal and db).
With the 14.2.10, I see on the bluestore OSDs a lot of read IOPs while snaptrimming. With 14.2.9 this was not an issue.
I wonder if this would explain the huge amount of slowops on my big testcluster (44 Nodes 1056 OSDs) while snaptrimming. I cannot test a downgrade there, because there are no packages of older releases for CentOS 8 available.
Regards Manuel
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
In my case I only have 16GB RAM per node with 5 OSD on each of them, so I actually have to tune osd_memory_target=2147483648 because with the default value of 4GB my osd processes tend to get killed by OOM. That is what I was looking into before the correct solution. I disabled osd_memory_target limitation essentially setting it to default 4GB - it helped in a sense that workload on the block.db device significantly dropped, but overall pattern was not the same - for example there still were no merges on the block.db device. It all came back to the usual pattern with bluefs_buffered_io=true. osd_memory_target limitation was implemented somewhere around 10 > 12 release upgrade I think, before memory auto scaling feature for bluestore was introduced - that's when my osds started to get OOM. They worked fine before that. чт, 6 авг. 2020 г. в 20:28, Mark Nelson <mnelson@redhat.com>:
Yeah, there are cases where enabling it will improve performance as rocksdb can then used the page cache as a (potentially large) secondary cache beyond the block cache and avoid hitting the underlying devices for reads. Do you have a lot of spare memory for page cache on your OSD nodes? You may be able to improve the situation with bluefs_buffered_io=false by increasing the osd_memory_target which should give the rocksdb block cache more memory to work with directly. One downside is that we currently double cache onodes in both the rocksdb cache and bluestore onode cache which hurts us when memory limited. We have some experimental work that might help in this area by better balancing bluestore onode and rocksdb block caches but it needs to be rebased after Adam's column family sharding work.
The reason we had to disable bluefs_buffered_io again was that we had users with certain RGW workloads where the kernel started swapping large amounts of memory on the OSD nodes despite seemingly have free memory available. This caused huge latency spikes and IO slowdowns (even stalls). We never noticed it in our QA test suites and it doesn't appear to happen with RBD workloads as far as I can tell, but when it does happen it's really painful.
Mark
On 8/6/20 6:53 AM, Manuel Lausch wrote:
Hi,
I found the reasen of this behavior change. With 14.2.10 the default value of "bluefs_buffered_io" was changed from true to false. https://tracker.ceph.com/issues/44818
configureing this to true my problems seems to be solved.
Regards Manuel
On Wed, 5 Aug 2020 13:30:45 +0200 Manuel Lausch <manuel.lausch@1und1.de> wrote:
Hello Vladimir,
I just tested this with a single node testcluster with 60 HDDs (3 of them with bluestore without separate wal and db).
With the 14.2.10, I see on the bluestore OSDs a lot of read IOPs while snaptrimming. With 14.2.9 this was not an issue.
I wonder if this would explain the huge amount of slowops on my big testcluster (44 Nodes 1056 OSDs) while snaptrimming. I cannot test a downgrade there, because there are no packages of older releases for CentOS 8 available.
Regards Manuel
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
a 2GB memory target will absolutely starve the OSDs of memory for rocksdb block cache which probably explains why you are hitting the disk for reads and a shared page cache is helping so much. It's definitely more memory efficient to have a page cache scheme rather than having more cache for each OSD, but for NVMe drives you can end up having more contention and overhead. For older systems with slower devices and lower amounts of memory the page cache is probably a win. FWIW with a 4GB+ memory target I suspect you would see far fewer cache miss reads (but obviously you can't do that on your nodes). Mark On 8/6/20 1:47 PM, Vladimir Prokofev wrote:
In my case I only have 16GB RAM per node with 5 OSD on each of them, so I actually have to tune osd_memory_target=2147483648 because with the default value of 4GB my osd processes tend to get killed by OOM. That is what I was looking into before the correct solution. I disabled osd_memory_target limitation essentially setting it to default 4GB - it helped in a sense that workload on the block.db device significantly dropped, but overall pattern was not the same - for example there still were no merges on the block.db device. It all came back to the usual pattern with bluefs_buffered_io=true. osd_memory_target limitation was implemented somewhere around 10 > 12 release upgrade I think, before memory auto scaling feature for bluestore was introduced - that's when my osds started to get OOM. They worked fine before that.
чт, 6 авг. 2020 г. в 20:28, Mark Nelson <mnelson@redhat.com>:
Yeah, there are cases where enabling it will improve performance as rocksdb can then used the page cache as a (potentially large) secondary cache beyond the block cache and avoid hitting the underlying devices for reads. Do you have a lot of spare memory for page cache on your OSD nodes? You may be able to improve the situation with bluefs_buffered_io=false by increasing the osd_memory_target which should give the rocksdb block cache more memory to work with directly. One downside is that we currently double cache onodes in both the rocksdb cache and bluestore onode cache which hurts us when memory limited. We have some experimental work that might help in this area by better balancing bluestore onode and rocksdb block caches but it needs to be rebased after Adam's column family sharding work.
The reason we had to disable bluefs_buffered_io again was that we had users with certain RGW workloads where the kernel started swapping large amounts of memory on the OSD nodes despite seemingly have free memory available. This caused huge latency spikes and IO slowdowns (even stalls). We never noticed it in our QA test suites and it doesn't appear to happen with RBD workloads as far as I can tell, but when it does happen it's really painful.
Mark
On 8/6/20 6:53 AM, Manuel Lausch wrote:
Hi,
I found the reasen of this behavior change. With 14.2.10 the default value of "bluefs_buffered_io" was changed from true to false. https://tracker.ceph.com/issues/44818
configureing this to true my problems seems to be solved.
Regards Manuel
On Wed, 5 Aug 2020 13:30:45 +0200 Manuel Lausch <manuel.lausch@1und1.de> wrote:
Hello Vladimir,
I just tested this with a single node testcluster with 60 HDDs (3 of them with bluestore without separate wal and db).
With the 14.2.10, I see on the bluestore OSDs a lot of read IOPs while snaptrimming. With 14.2.9 this was not an issue.
I wonder if this would explain the huge amount of slowops on my big testcluster (44 Nodes 1056 OSDs) while snaptrimming. I cannot test a downgrade there, because there are no packages of older releases for CentOS 8 available.
Regards Manuel
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I cannot confirm that more memory target will solve the problem completly. In my case the OSDs have 14GB memory target and I did have huge user IO impact while snaptrim (many slow ops the whole time). Since I set bluefs_bufferd_io=true it seems to work without issue. In my cluster I don't use rgw. But I don't see why different types of access the cluster do affect the form the kernel manages its memory. My experience why the kernel begins to swap are mostly numa related and/or memory fragmentation. Manuel On Thu, 6 Aug 2020 15:06:49 -0500 Mark Nelson <mnelson@redhat.com> wrote:
a 2GB memory target will absolutely starve the OSDs of memory for rocksdb block cache which probably explains why you are hitting the disk for reads and a shared page cache is helping so much. It's definitely more memory efficient to have a page cache scheme rather than having more cache for each OSD, but for NVMe drives you can end up having more contention and overhead. For older systems with slower devices and lower amounts of memory the page cache is probably a win. FWIW with a 4GB+ memory target I suspect you would see far fewer cache miss reads (but obviously you can't do that on your nodes).
Mark
On 8/6/20 1:47 PM, Vladimir Prokofev wrote:
In my case I only have 16GB RAM per node with 5 OSD on each of them, so I actually have to tune osd_memory_target=2147483648 because with the default value of 4GB my osd processes tend to get killed by OOM. That is what I was looking into before the correct solution. I disabled osd_memory_target limitation essentially setting it to default 4GB - it helped in a sense that workload on the block.db device significantly dropped, but overall pattern was not the same - for example there still were no merges on the block.db device. It all came back to the usual pattern with bluefs_buffered_io=true. osd_memory_target limitation was implemented somewhere around 10 > 12 release upgrade I think, before memory auto scaling feature for bluestore was introduced - that's when my osds started to get OOM. They worked fine before that.
чт, 6 авг. 2020 г. в 20:28, Mark Nelson <mnelson@redhat.com>:
Yeah, there are cases where enabling it will improve performance as rocksdb can then used the page cache as a (potentially large) secondary cache beyond the block cache and avoid hitting the underlying devices for reads. Do you have a lot of spare memory for page cache on your OSD nodes? You may be able to improve the situation with bluefs_buffered_io=false by increasing the osd_memory_target which should give the rocksdb block cache more memory to work with directly. One downside is that we currently double cache onodes in both the rocksdb cache and bluestore onode cache which hurts us when memory limited. We have some experimental work that might help in this area by better balancing bluestore onode and rocksdb block caches but it needs to be rebased after Adam's column family sharding work.
The reason we had to disable bluefs_buffered_io again was that we had users with certain RGW workloads where the kernel started swapping large amounts of memory on the OSD nodes despite seemingly have free memory available. This caused huge latency spikes and IO slowdowns (even stalls). We never noticed it in our QA test suites and it doesn't appear to happen with RBD workloads as far as I can tell, but when it does happen it's really painful.
Mark
On 2020-08-07 09:27, Manuel Lausch wrote:
I cannot confirm that more memory target will solve the problem completly. In my case the OSDs have 14GB memory target and I did have huge user IO impact while snaptrim (many slow ops the whole time). Since I set bluefs_bufferd_io=true it seems to work without issue. In my cluster I don't use rgw. But I don't see why different types of access the cluster do affect the form the kernel manages its memory. My experience why the kernel begins to swap are mostly numa related and/or memory fragmentation.
Can you share the amount of buffer cache available on your storage nodes? We run the OSDs with osd_memory_target=11G and 22 GB of buffer cache available. And with the buffer on (Mimic 13.2.8). Thanks, Gr. Stefan -- | BIT BV https://www.bit.nl/ Kamer van Koophandel 09090351 | GPG: 0xD14839C6 +31 318 648 688 / info@bit.nl
Sure. total used free shared buff/cache available Mem: 394582604 355494400 5282932 1784 33805272 29187220 Swap: 1047548 1047548 0 On the node are 24 14TB OSDs with 14G configured memory target. Manuel On Fri, 7 Aug 2020 11:24:53 +0200 Stefan Kooman <stefan@bit.nl> wrote:
Can you share the amount of buffer cache available on your storage nodes?
We run the OSDs with osd_memory_target=11G and 22 GB of buffer cache available. And with the buffer on (Mimic 13.2.8).
Thanks,
Gr. Stefan
It's quite possible that the issue is really about rocksdb living on top of bluefs with bluefs_buffered_io and rgw causing a ton of OMAP traffic. rgw is the only case so far where the issue has shown up, but it was significant enough that we didn't feel like we could leave bluefs_buffered_io enabled. In your case with a 14GB target per OSD, do you still see significantly increased disk reads with blufs_buffered_io=false? Mark On 8/7/20 2:27 AM, Manuel Lausch wrote:
I cannot confirm that more memory target will solve the problem completly. In my case the OSDs have 14GB memory target and I did have huge user IO impact while snaptrim (many slow ops the whole time). Since I set bluefs_bufferd_io=true it seems to work without issue. In my cluster I don't use rgw. But I don't see why different types of access the cluster do affect the form the kernel manages its memory. My experience why the kernel begins to swap are mostly numa related and/or memory fragmentation.
Manuel
On Thu, 6 Aug 2020 15:06:49 -0500 Mark Nelson <mnelson@redhat.com> wrote:
a 2GB memory target will absolutely starve the OSDs of memory for rocksdb block cache which probably explains why you are hitting the disk for reads and a shared page cache is helping so much. It's definitely more memory efficient to have a page cache scheme rather than having more cache for each OSD, but for NVMe drives you can end up having more contention and overhead. For older systems with slower devices and lower amounts of memory the page cache is probably a win. FWIW with a 4GB+ memory target I suspect you would see far fewer cache miss reads (but obviously you can't do that on your nodes).
Mark
On 8/6/20 1:47 PM, Vladimir Prokofev wrote:
In my case I only have 16GB RAM per node with 5 OSD on each of them, so I actually have to tune osd_memory_target=2147483648 because with the default value of 4GB my osd processes tend to get killed by OOM. That is what I was looking into before the correct solution. I disabled osd_memory_target limitation essentially setting it to default 4GB - it helped in a sense that workload on the block.db device significantly dropped, but overall pattern was not the same - for example there still were no merges on the block.db device. It all came back to the usual pattern with bluefs_buffered_io=true. osd_memory_target limitation was implemented somewhere around 10 > 12 release upgrade I think, before memory auto scaling feature for bluestore was introduced - that's when my osds started to get OOM. They worked fine before that.
чт, 6 авг. 2020 г. в 20:28, Mark Nelson <mnelson@redhat.com>:
Yeah, there are cases where enabling it will improve performance as rocksdb can then used the page cache as a (potentially large) secondary cache beyond the block cache and avoid hitting the underlying devices for reads. Do you have a lot of spare memory for page cache on your OSD nodes? You may be able to improve the situation with bluefs_buffered_io=false by increasing the osd_memory_target which should give the rocksdb block cache more memory to work with directly. One downside is that we currently double cache onodes in both the rocksdb cache and bluestore onode cache which hurts us when memory limited. We have some experimental work that might help in this area by better balancing bluestore onode and rocksdb block caches but it needs to be rebased after Adam's column family sharding work.
The reason we had to disable bluefs_buffered_io again was that we had users with certain RGW workloads where the kernel started swapping large amounts of memory on the OSD nodes despite seemingly have free memory available. This caused huge latency spikes and IO slowdowns (even stalls). We never noticed it in our QA test suites and it doesn't appear to happen with RBD workloads as far as I can tell, but when it does happen it's really painful.
Mark
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Mark, The read IOPs in "normal" operation was with bluefs_buffered_io=false somewhat about 1. And now with true around 2. So this seems slightly higher, but far away from any problem. While snapshot trimming the difference is enormous. with false: around 200 with true: around 10 scrubing read IOPs do not appear to be affected. They are around 100 IOPs I'am using librados to access my objects. So I don't know if this would be any different with rgw. Manuel On Fri, 7 Aug 2020 08:08:40 -0500 Mark Nelson <mnelson@redhat.com> wrote:
It's quite possible that the issue is really about rocksdb living on top of bluefs with bluefs_buffered_io and rgw causing a ton of OMAP traffic. rgw is the only case so far where the issue has shown up, but it was significant enough that we didn't feel like we could leave bluefs_buffered_io enabled. In your case with a 14GB target per OSD, do you still see significantly increased disk reads with blufs_buffered_io=false?
Mark
That is super interesting regarding scrubbing. I would have expected that to be affected as well. Any chance you can check and see if there is any correlation between rocksdb compaction events, snap trimming, and increased disk reads? Also (Sorry if you already answered this) do we know for sure that it's hitting the block.db/block.wal device? I suspect it is, just wanted to verify. Mark On 8/7/20 9:04 AM, Manuel Lausch wrote:
Hi Mark,
The read IOPs in "normal" operation was with bluefs_buffered_io=false somewhat about 1. And now with true around 2. So this seems slightly higher, but far away from any problem.
While snapshot trimming the difference is enormous. with false: around 200 with true: around 10
scrubing read IOPs do not appear to be affected. They are around 100 IOPs
I'am using librados to access my objects. So I don't know if this would be any different with rgw.
Manuel
On Fri, 7 Aug 2020 08:08:40 -0500 Mark Nelson <mnelson@redhat.com> wrote:
It's quite possible that the issue is really about rocksdb living on top of bluefs with bluefs_buffered_io and rgw causing a ton of OMAP traffic. rgw is the only case so far where the issue has shown up, but it was significant enough that we didn't feel like we could leave bluefs_buffered_io enabled. In your case with a 14GB target per OSD, do you still see significantly increased disk reads with blufs_buffered_io=false?
Mark
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Mark, rocskdb compactions was one of my first ideas as well. But they don't correlate. I checkt this with the ceph_rocskdb_log_parser.py from https://github.com/ceph/cbt.git I saw only a few compactions on the whole cluster. It didn't seem to be the problem, although the compactions sometimes took several seconds. BTW: I configured the following rocksdb options. bluestore rocksdb options = compression=kNoCompression,max_write_buffer_number=32,min_write_buffer_number_to_merge=2,recycle_log_file_num=32,compaction_style=kCompactionStyleLevel,write_buffer_size=67108864,target_file_size_base=67108864,max_background_compactions=31,level0_file_num_compaction_trigger=8,level0_slowdown_writes_trigger=32,level0_stop_writes_trigger=64,max_bytes_for_level_base=536870912,compaction_threads=32,max_bytes_for_level_multiplier=8,flusher_threads=8,compaction_readahead_size=2MB This reduced some IO spikes but the slowops isse while snaptim was not affected by this. Manuel On Fri, 7 Aug 2020 09:43:51 -0500 Mark Nelson <mnelson@redhat.com> wrote:
That is super interesting regarding scrubbing. I would have expected that to be affected as well. Any chance you can check and see if there is any correlation between rocksdb compaction events, snap trimming, and increased disk reads? Also (Sorry if you already answered this) do we know for sure that it's hitting the block.db/block.wal device? I suspect it is, just wanted to verify.
Mark
Yeah, I know various folks have adopted those settings, though I'm not convinced they are better than our defaults. Basically you have more smaller buffers and start compacting sooner and theoretically should have a more gradual throttle along with a bunch of changes to compaction, but every time I've tried a setup like that I see more write amplification in L0 presumably due to a larger number of pglog entries not being tomstoned before hitting it (at least on our systems it's not faster at this time, and imposes more wear on DB device). I suspect something closer to those settings will be better though if we can change the pglog to create/delete new kv pairs for every pglog entry. In any event, that's good to know about compaction not being involved. I think this may be a case where the double-caching fix might help significantly if we stop thrashing the rocksdb block cache: https://github.com/ceph/ceph/pull/27705 Mark On 8/10/20 2:28 AM, Manuel Lausch wrote:
Hi Mark,
rocskdb compactions was one of my first ideas as well. But they don't correlate. I checkt this with the ceph_rocskdb_log_parser.py from https://github.com/ceph/cbt.git I saw only a few compactions on the whole cluster. It didn't seem to be the problem, although the compactions sometimes took several seconds.
BTW: I configured the following rocksdb options. bluestore rocksdb options = compression=kNoCompression,max_write_buffer_number=32,min_write_buffer_number_to_merge=2,recycle_log_file_num=32,compaction_style=kCompactionStyleLevel,write_buffer_size=67108864,target_file_size_base=67108864,max_background_compactions=31,level0_file_num_compaction_trigger=8,level0_slowdown_writes_trigger=32,level0_stop_writes_trigger=64,max_bytes_for_level_base=536870912,compaction_threads=32,max_bytes_for_level_multiplier=8,flusher_threads=8,compaction_readahead_size=2MB
This reduced some IO spikes but the slowops isse while snaptim was not affected by this.
Manuel
On Fri, 7 Aug 2020 09:43:51 -0500 Mark Nelson <mnelson@redhat.com> wrote:
That is super interesting regarding scrubbing. I would have expected that to be affected as well. Any chance you can check and see if there is any correlation between rocksdb compaction events, snap trimming, and increased disk reads? Also (Sorry if you already answered this) do we know for sure that it's hitting the block.db/block.wal device? I suspect it is, just wanted to verify.
Mark
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi, I'm facing this issue too and I see the attached rocksdb log from Mark in my cluster which means there is a burst read on my block.db. I've sent some information from my issue in this thread[1]. Hope you help me with what's going on in my cluster. Thanks. [1]: https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/PHB53F3OD7QN... On Mon, Aug 10, 2020 at 8:05 PM Mark Nelson <mnelson@redhat.com> wrote:
Yeah, I know various folks have adopted those settings, though I'm not convinced they are better than our defaults. Basically you have more smaller buffers and start compacting sooner and theoretically should have a more gradual throttle along with a bunch of changes to compaction, but every time I've tried a setup like that I see more write amplification in L0 presumably due to a larger number of pglog entries not being tomstoned before hitting it (at least on our systems it's not faster at this time, and imposes more wear on DB device). I suspect something closer to those settings will be better though if we can change the pglog to create/delete new kv pairs for every pglog entry.
In any event, that's good to know about compaction not being involved. I think this may be a case where the double-caching fix might help significantly if we stop thrashing the rocksdb block cache: https://github.com/ceph/ceph/pull/27705
Mark
On 8/10/20 2:28 AM, Manuel Lausch wrote:
Hi Mark,
rocskdb compactions was one of my first ideas as well. But they don't correlate. I checkt this with the ceph_rocskdb_log_parser.py from https://github.com/ceph/cbt.git I saw only a few compactions on the whole cluster. It didn't seem to be the problem, although the compactions sometimes took several seconds.
BTW: I configured the following rocksdb options. bluestore rocksdb options = compression=kNoCompression,max_write_buffer_number=32,min_write_buffer_number_to_merge=2,recycle_log_file_num=32,compaction_style=kCompactionStyleLevel,write_buffer_size=67108864,target_file_size_base=67108864,max_background_compactions=31,level0_file_num_compaction_trigger=8,level0_slowdown_writes_trigger=32,level0_stop_writes_trigger=64,max_bytes_for_level_base=536870912,compaction_threads=32,max_bytes_for_level_multiplier=8,flusher_threads=8,compaction_readahead_size=2MB
This reduced some IO spikes but the slowops isse while snaptim was not affected by this.
Manuel
On Fri, 7 Aug 2020 09:43:51 -0500 Mark Nelson <mnelson@redhat.com> wrote:
That is super interesting regarding scrubbing. I would have expected that to be affected as well. Any chance you can check and see if there is any correlation between rocksdb compaction events, snap trimming, and increased disk reads? Also (Sorry if you already answered this) do we know for sure that it's hitting the block.db/block.wal device? I suspect it is, just wanted to verify.
Mark
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (6)
-
Eneko Lacunza
-
Manuel Lausch
-
Mark Nelson
-
Seena Fallah
-
Stefan Kooman
-
Vladimir Prokofev