Re: BlueStore not surviving power outage
Thanks, everyone. There is a RAID HBA in each of the machines in our clusters, to which all SATA disks are attached. We configured the RAID HBA cache mode to "write through", but, as I checked yesterday, the BBU of the RAID HBAs are not charged. I'm not quite sure whether the BBU has something to do with the data loss, as far as I know, all data should be persisted to the underlying disk before acknowledging upper layer systems when cache mode is "write through". Am I missing anything? Thanks:-) On Tue, 27 Apr 2021 at 23:43, Martin Verges <martin.verges@croit.io> wrote:
What drives do you use? Do they have PLP (power loss protection)? Is there any form of raid controller involved?
-- Martin Verges Managing director
Mobile: +49 174 9335695 E-Mail: martin.verges@croit.io Chat: https://t.me/MartinVerges
croit GmbH, Freseniusstr. 31h, 81247 Munich CEO: Martin Verges - VAT-ID: DE310638492 Com. register: Amtsgericht Munich HRB 231263
Web: https://croit.io YouTube: https://goo.gl/PGE1Bx
On Tue, 27 Apr 2021 at 10:54, Xuehan Xu <xxhdx1985126@gmail.com> wrote:
Hi, everyone.
Recently, one of our online cluster experienced a whole cluster power outage, and after the power recovered, many osd started to log the following error:
2021-04-27 15:38:05.503 2b372b957700 -1 bluestore(/var/lib/ceph/osd/ceph-3) _verify_csum bad crc32c/0x1000 checksum at blob offset 0x36000, got 0x41fe1397, expected 0x8d7f5975, device location [0xa7e76000~1000], logical extent 0x1b6000~1000, object #9:45a4e02a:::rbd_data.3b35df93038d.0000000000000095:head# 2021-04-27 15:38:05.504 2b372b957700 -1 bluestore(/var/lib/ceph/osd/ceph-3) _verify_csum bad crc32c/0x1000 checksum at blob offset 0x36000, got 0x41fe1397, expected 0x8d7f5975, device location [0xa7e76000~1000], logical extent 0x1b6000~1000, object #9:45a4e02a:::rbd_data.3b35df93038d.0000000000000095:head# 2021-04-27 15:38:05.505 2b372b957700 -1 bluestore(/var/lib/ceph/osd/ceph-3) _verify_csum bad crc32c/0x1000 checksum at blob offset 0x36000, got 0x41fe1397, expected 0x8d7f5975, device location [0xa7e76000~1000], logical extent 0x1b6000~1000, object #9:45a4e02a:::rbd_data.3b35df93038d.0000000000000095:head# 2021-04-27 15:38:05.506 2b372b957700 -1 bluestore(/var/lib/ceph/osd/ceph-3) _verify_csum bad crc32c/0x1000 checksum at blob offset 0x36000, got 0x41fe1397, expected 0x8d7f5975, device location [0xa7e76000~1000], logical extent 0x1b6000~1000, object #9:45a4e02a:::rbd_data.3b35df93038d.0000000000000095:head# 2021-04-27 15:38:28.379 2b372c158700 -1 bluestore(/var/lib/ceph/osd/ceph-3) _verify_csum bad crc32c/0x1000 checksum at blob offset 0x40000, got 0xce935e16, expected 0x9b502da7, device location [0xa9a80000~1000], logical extent 0x80000~1000, object #9:c2a6d9ae:::rbd_data.3b35df93038d.0000000000000696:head#
We are using Nautilus 14.2.10 version, and we put rocksdb on top of SSDs while bluestore data on SATA disks. It seems that the BlueStore didn't survive the power outage, is it supposed to behave this way? Is there any way to prevent it?
Thanks:-) _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Den ons 28 apr. 2021 kl 04:25 skrev Xuehan Xu <xxhdx1985126@gmail.com>:
There is a RAID HBA in each of the machines in our clusters, to which all SATA disks are attached. We configured the RAID HBA cache mode to "write through", but, as I checked yesterday, the BBU of the RAID HBAs are not charged. I'm not quite sure whether the BBU has something to do with the data loss, as far as I know, all data should be persisted to the underlying disk before acknowledging upper layer systems when cache mode is "write through". Am I missing anything? Thanks:-)
If the raid card was good, it would change caching strategy when/if the BBU has no power left, but if it didn't and "it was good when we last booted up", then it is possible that it 'promised' that writing to BBU-backed RAM was ok for ack'ing the writes even if they are not on disk yet, and when the BBU failed (for whatever reason), then this promise was not honored and lots of writes were lost. -- May the most significant bit of your life be positive.
On 28/04/2021 09:13, Janne Johansson wrote:
Den ons 28 apr. 2021 kl 04:25 skrev Xuehan Xu <xxhdx1985126@gmail.com>:
There is a RAID HBA in each of the machines in our clusters, to which all SATA disks are attached. We configured the RAID HBA cache mode to "write through", but, as I checked yesterday, the BBU of the RAID HBAs are not charged. I'm not quite sure whether the BBU has something to do with the data loss, as far as I know, all data should be persisted to the underlying disk before acknowledging upper layer systems when cache mode is "write through". Am I missing anything? Thanks:-)
If the raid card was good, it would change caching strategy when/if the BBU has no power left, but if it didn't and "it was good when we last booted up", then it is possible that it 'promised' that writing to BBU-backed RAM was ok for ack'ing the writes even if they are not on disk yet, and when the BBU failed (for whatever reason), then this promise was not honored and lots of writes were lost.
In addition: In 2020 I have seen two cases (as a Ceph consultant) of severe data corruption with BlueStore after a power failure. In both cases this happened on systems where an HBA was involved. In the end we blaimed the HBAs which were in RAID mode. I have done extensive power failure testing afterwards on NVMe-only and on systems with HBAs in JBOD mode and I was never able to reproduce the data corruption after a power failure. My suspicion is still that the HBAs were caching some data and it was not written to the medium before the power failed although BlueStore was told it was. My bet: This is the HBA, not BlueStore's fault. Wido
If the raid card was good, it would change caching strategy when/if the BBU has no power left, but if it didn't and "it was good when we last booted up", then it is possible that it 'promised' that writing to BBU-backed RAM was ok for ack'ing the writes even if they are not on disk yet, and when the BBU failed (for whatever reason), then this promise was not honored and lots of writes were lost.
In addition: In 2020 I have seen two cases (as a Ceph consultant) of severe data corruption with BlueStore after a power failure.
In both cases this happened on systems where an HBA was involved. In the end we blaimed the HBAs which were in RAID mode.
I have done extensive power failure testing afterwards on NVMe-only and on systems with HBAs in JBOD mode and I was never able to reproduce the data corruption after a power failure.
My suspicion is still that the HBAs were caching some data and it was not written to the medium before the power failed although BlueStore was told it was.
My bet: This is the HBA, not BlueStore's fault.
Wido
I fully agree. cf. a post I made … nearly two years ago. http://lists.ceph.com/pipermail/ceph-users-ceph.com/2019-July/036237.html I had all manner of similar experiences with RoC HBAs: - cache retention and flushing bugs - hardware issues requiring cards (300+ just in my dept) to be reworked that — you guessed it — prevented cache preservation from working properly - firmware and management utility that silently enabled the drives’ volatile cache — and lied about it. - Flaky BBU/supercap modules with rather finicky connectors. I would not be surprised if the HBAs in question are a few years behind in firmware updates. The wrap-every-drive-in-a-VD strategy is all too familiar. Setting the HBA into the JBOD personality, enabling passthrough, or flashing with IT firmware are what I suggest, depending on the model in question.
participants (4)
-
Anthony D'Atri
-
Janne Johansson
-
Wido den Hollander
-
Xuehan Xu