Hello Ceph Users, * Problem: we get the following errors when using krbd, we are using rbd for vms. * Workaround: by switching to librbd the errors disappear. * Software: ** Kernel: 6.8.8-2 (parameters: intel_iommu=on iommu=pt pcie_aspm.policy=performance) ** Ceph: 18.2.2 Description/Details: Errors from using krbd with ceph. Side-effects: [Wed Aug 21 03:04:17 2024] libceph: read_partial_message 0000000015af2284 data crc 1221767919 != exp. 282251377 [Wed Aug 21 03:04:17 2024] libceph: read_partial_message 0000000066b200ab data crc 3817026135 != exp. 3925662391 [Wed Aug 21 03:04:17 2024] libceph: osd15 (1)10.1.4.13:6836 bad crc/signature [Wed Aug 21 03:04:17 2024] libceph: osd13 (1)10.1.4.13:6809 bad crc/signature [Wed Aug 21 03:04:21 2024] libceph: read_partial_message 000000008a131738 data crc 2612835980 != exp. 917302924 [Wed Aug 21 03:04:21 2024] libceph: read_partial_message 000000005160776b data crc 2965872045 != exp. 563323792 [Wed Aug 21 03:04:21 2024] libceph: osd15 (1)10.1.4.13:6836 bad crc/signature [Wed Aug 21 03:04:21 2024] libceph: osd6 (1)10.1.4.12:6843 bad crc/signature [Wed Aug 21 03:06:44 2024] libceph: read_partial_message 000000007e548354 data crc 1265032637 != exp. 2426281931 [Wed Aug 21 03:06:44 2024] libceph: osd0 (1)10.1.4.11:6835 bad crc/signature [Wed Aug 21 03:06:44 2024] libceph: read_partial_message 000000009214d802 data crc 2596010853 != exp. 1221875667 [Wed Aug 21 03:06:44 2024] libceph: osd10 (1)10.1.4.12:6809 bad crc/signature [Wed Aug 21 03:06:47 2024] libceph: read_partial_message 000000000f9edc73 data crc 1326019705 != exp. 3079604517 [Wed Aug 21 03:06:47 2024] libceph: osd3 (1)10.1.4.11:6803 bad crc/signature [Wed Aug 21 03:06:50 2024] libceph: read_partial_message 000000004769da61 data crc 3421275194 != exp. 4183754554 [Wed Aug 21 03:06:50 2024] libceph: osd8 (1)10.1.4.12:6835 bad crc/signature [Wed Aug 21 03:06:51 2024] libceph: read_partial_message 0000000044db9a59 data crc 2603270708 != exp. 4150529351 [Wed Aug 21 03:06:51 2024] libceph: osd14 (1)10.1.4.13:6806 bad crc/signature Description/Details 2: vms get problems with buffer i/o errors on rbd-backed virtual disks: Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#211 timing out command, waited 180s Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#39 timing out command, waited 180s Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#211 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=885s Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#39 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=847s Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#211 Sense Key: Aborted Command [current] Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#211 Add. Sense: I/O process terminated Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#39 Sense Key: Aborted Command [current] Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#39 Add. Sense: I/0 process terminated Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#211 CDB: Write(10) 2a 00 34 87 48 08 00 00 08 00 Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#39 CDB: Write (10) 2a 00 34 81 52 70 00 00 58 00 Aug 24 03:16:01 de-vlix-dbix-01 kernel: I/O error, dev sdb, sector 881281032 op 0x1: (WRITE) flags 0x800 phys_seg 1 prio class 0 Aug 24 03:16:01 de-vlix-dbix-01 kernel: I/O error, dev sdb, sector 880890480 op 0x1: (WRITE) flags 0x103000 phys_seg 11 prio class 0 Aug 24 03:16:01 de-vlix-dbix-01 kernel: EXT4-fs warning (device sdb1): ext4_end_bio:343: I/0 error 10 writing to inode 27525908 starting Aug 24 03:16:01 de-vlix-dbix-01 kernel: Buffer I/O error on dev sdbl, logical block 110111054, lost async page write Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#212 timing out command, waited 180s Aug 24 03:16:01 de-vlix-dbix-01 kernel: buffer_io_error: 21 callbacks suppressed Aug 24 03:16:01 de-vlix-dbix-01 kernel: Buffer 1/0 error on device sdbl, logical block 110159873 Aug 24 03:16:01 de-vlix-dbix-01 kernel: Buffer 1/0 error on dev sdbl, logical block 110111055, lost async page write Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#212 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_OK cmd_age=875s Aug 24 03:16:01 de-vlix-dbix-01 kernel: Buffer I/O error on dev sdb1, logical block 110111056, lost async page write Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#212 Sense Key: Aborted Command [current] Aug 24 03:16:01 de-vlix-dbix-01 kernel: Buffer I/O error on dev sdbl, logical block 110111057, lost async page write Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#212 Add. Sense: I/0 process terminated Aug 24 03:16:01 de-vlix-dbix-01 kernel: Buffer I/O error on dev sdb1, logical block 110111058, lost async page write Aug 24 03:16:01 de-vlix-dbix-01 kernel: sd 0:0:0:1: [sdb] tag#212 CDB: Write(10) 2a 00 34 87 48 10 00 00 08 00 Aug 24 03:16:01 de-vlix-dbix-01 kernel: Buffer I/O error on dev sdbl, logical block 110111059, lost async page write Aug 24 03:16:01 de-vlix-dbix-01 kernel: I/O error, dev sdb, sector 881281040 op 0x1: (WRITE) flags 0x800 phys_seg 1 prio class 0 Aug 24 03:16:01 de-vlix-dbix-01 kernel: Buffer I/O error on dev sdbl, logical block 110111060, lost async page write Thanks for helping out. Greetings Jonas
On Fri, Sep 6, 2024 at 3:54 AM <jsterr@deckplosion.de> wrote:
Hello Ceph Users,
* Problem: we get the following errors when using krbd, we are using rbd for vms. * Workaround: by switching to librbd the errors disappear.
* Software: ** Kernel: 6.8.8-2 (parameters: intel_iommu=on iommu=pt pcie_aspm.policy=performance) ** Ceph: 18.2.2
Description/Details: Errors from using krbd with ceph. Side-effects:
[Wed Aug 21 03:04:17 2024] libceph: read_partial_message 0000000015af2284 data crc 1221767919 != exp. 282251377 [Wed Aug 21 03:04:17 2024] libceph: read_partial_message 0000000066b200ab data crc 3817026135 != exp. 3925662391 [Wed Aug 21 03:04:17 2024] libceph: osd15 (1)10.1.4.13:6836 bad crc/signature [Wed Aug 21 03:04:17 2024] libceph: osd13 (1)10.1.4.13:6809 bad crc/signature
Hi Jonas, Are these VMs running Windows? If so, are you using rxbounce mapping option ("rbd device map -o rxbounce ...")? It's more or less required in case there is a Windows kernel on the I/O path: rxbounce - Use a bounce buffer when receiving data (since 5.17). The default behaviour is to read directly into the destination buffer. A bounce buffer is needed if the destination buffer isn’t guaranteed to be stable (i.e. remain unchanged while it is being read to). In particular this is the case for Windows where a system-wide “dummy” (throwaway) page may be mapped into the destination buffer in order to generate a single large I/O. Otherwise, “libceph: … bad crc/signature” or “libceph: … integrity error, bad crc” errors and associated performance degradation are expected. Thanks, Ilya
Hi Ilya, some of our Proxmox VE users also report they need to enable rxbounce to avoid their Windows VMs triggering these errors, see e.g. [1]. With rxbounce, everything seems to work smoothly, so thanks for adding this option. :) We're currently checking how our stack could handle this more gracefully. From my understanding of the rxbounce option, it seems like always passing it when mapping a volume (i.e., even if the VM disks belong to Linux VMs that aren't affected by this issue) might not be a good idea performance-wise. Another option is to only pass `rxbounce` when mapping volumes that are known to be Windows VM disks. This seems like the most sensible option, but still, I wanted to ask: I see that when rxbounce was originally introduced [2], the possibility of automatically switching over to "rxbounce mode" when needed was also discussed. Do you think this is a direction that krbd might take some time in the future? Best, Friedrich [1] https://forum.proxmox.com/threads/155741/ [2] https://lore.kernel.org/all/894a36483c241e0cc5154e09e8dd078f57a606d5.camel@k...
On Thu, Dec 12, 2024 at 5:37 PM Friedrich Weber <f.weber@proxmox.com> wrote:
Hi Ilya,
some of our Proxmox VE users also report they need to enable rxbounce to avoid their Windows VMs triggering these errors, see e.g. [1]. With rxbounce, everything seems to work smoothly, so thanks for adding this option. :)
We're currently checking how our stack could handle this more gracefully. From my understanding of the rxbounce option, it seems like always passing it when mapping a volume (i.e., even if the VM disks belong to Linux VMs that aren't affected by this issue) might not be a good idea performance-wise.
Hi Friedrich, Yup, enabling rxbounce causes the read data to be double-buffered. However, I don't have any numbers on hand. The overhead may turn out to be negligible in your case.
Another option is to only pass `rxbounce` when mapping volumes that are known to be Windows VM disks. This seems like the most sensible option,
*nod*
but still, I wanted to ask: I see that when rxbounce was originally introduced [2], the possibility of automatically switching over to "rxbounce mode" when needed was also discussed. Do you think this is a direction that krbd might take some time in the future?
Very unlikely. Trying to enable rxbounce behind the scenes after observing some CRC errors bumps into questions like what should be the threshold in terms of the number of errors and also the number of OSD sessions that are affected, should rxbounce be enabled globally or just for those OSD sessions, should it be disabled after time passes, etc. Fundamentally, krbd can't distinguish between "legit" CRC errors and something that just needs to be worked around. There is a concern that enabling rxbounce automatically could mask bugs or behaviors similar to what the Windows kernel does with its dummy page which we (developers) would like to know about ;) Thanks, Ilya
Hi Ilya, On 13/12/2024 16:25, Ilya Dryomov wrote:
[...]
We're currently checking how our stack could handle this more gracefully. From my understanding of the rxbounce option, it seems like always passing it when mapping a volume (i.e., even if the VM disks belong to Linux VMs that aren't affected by this issue) might not be a good idea performance-wise.
Hi Friedrich,
Yup, enabling rxbounce causes the read data to be double-buffered. However, I don't have any numbers on hand. The overhead may turn out to be negligible in your case.
I see. I also ran some very quick benchmarks where I didn't see much impact, but that's not quite enough to be confident.
[...]
but still, I wanted to ask: I see that when rxbounce was originally introduced [2], the possibility of automatically switching over to "rxbounce mode" when needed was also discussed. Do you think this is a direction that krbd might take some time in the future?
Very unlikely. Trying to enable rxbounce behind the scenes after observing some CRC errors bumps into questions like what should be the threshold in terms of the number of errors and also the number of OSD sessions that are affected, should rxbounce be enabled globally or just for those OSD sessions, should it be disabled after time passes, etc. Fundamentally, krbd can't distinguish between "legit" CRC errors and something that just needs to be worked around. There is a concern that enabling rxbounce automatically could mask bugs or behaviors similar to what the Windows kernel does with its dummy page which we (developers) would like to know about ;)
That makes sense, thank you for elaborating! Best wishes, Friedrich
participants (3)
-
Friedrich Weber
-
Ilya Dryomov
-
jsterr@deckplosion.de