ceph iscsi latency too high for esxi?
Hi, does anyone here use CEPH iSCSI with VMware ESXi? It seems that we are hitting the 5 second timeout limit on software HBA in ESXi. It appears whenever there is increased load on the cluster, like deep scrub or rebalance. Is it normal behaviour in production? Or is there something special we need to tune? We are on latest Nautilus, 12 x 10 TB OSDs (4 servers), 25 Gbit/s Ethernet, erasure coded rbd pool with 128 PGs, aroun 200 PGs per OSD total. ESXi Log: 2020-10-04T01:57:04.314Z cpu34:2098959)WARNING: iscsi_vmk: iscsivmk_ConnReceiveAtomic:517: vmhba64:CH:1 T:0 CN:0: Failed to receive data: Connection closed by peer 2020-10-04T01:57:04.314Z cpu34:2098959)iscsi_vmk: iscsivmk_ConnRxNotifyFailure:1235: vmhba64:CH:1 T:0 CN:0: Connection rx notifying failure: Failed to Receive. State=Bound 2020-10-04T01:57:04.566Z cpu19:2098979)WARNING: iscsi_vmk: iscsivmk_StopConnection:741: vmhba64:CH:1 T:0 CN:0: iSCSI connection is being marked "OFFLINE" (Event:4) 2020-10-04T01:57:04.654Z cpu7:2097866)WARNING: VMW_SATP_ALUA: satp_alua_issueCommandOnPath:788: Probe cmd 0xa3 failed for path "vmhba64:C2:T0:L0" (0x5/0x20/0x0). Check if failover mode is still ALUA. OSD Log: [303088.450088] Did not receive response to NOPIN on CID: 0, failing connection for I_T Nexus iqn.1994-05.com.redhat:esxi1,i,0x00023d000002,iqn.2003-01.com.redhat.iscsi-gw:iscsi-igw,t,0x01 [324926.694077] Did not receive response to NOPIN on CID: 0, failing connection for I_T Nexus iqn.1994-05.com.redhat:esxi2,i,0x00023d000001,iqn.2003-01.com.redhat.iscsi-gw:iscsi-igw,t,0x01 [407067.404538] ABORT_TASK: Found referenced iSCSI task_tag: 5891 [407076.077175] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 5891 [411677.887690] ABORT_TASK: Found referenced iSCSI task_tag: 6722 [411683.297425] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 6722 [481459.755876] ABORT_TASK: Found referenced iSCSI task_tag: 7930 [481460.787968] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 7930 Cheers, Martin
Hello, no iSCSI + VMware works without such problems.
We are on latest Nautilus, 12 x 10 TB OSDs (4 servers), 25 Gbit/s Ethernet, erasure coded rbd pool with 128 PGs, aroun 200 PGs per OSD total.
Nautilus is a good choice 12*10TB HDD is not good for VMs 25Gbit/s on HDD is way to much for that system 200 PGs per OSD is to much, I would suggest 75-100 PGs per OSD You can improve latency on HDD clusters using external DB/WAL on NVMe. That might help you -- Martin Verges Managing director Mobile: +49 174 9335695 E-Mail: martin.verges@croit.io Chat: https://t.me/MartinVerges croit GmbH, Freseniusstr. 31h, 81247 Munich CEO: Martin Verges - VAT-ID: DE310638492 Com. register: Amtsgericht Munich HRB 231263 Web: https://croit.io YouTube: https://goo.gl/PGE1Bx Am So., 4. Okt. 2020 um 14:37 Uhr schrieb Golasowski Martin < martin.golasowski@vsb.cz>:
Hi, does anyone here use CEPH iSCSI with VMware ESXi? It seems that we are hitting the 5 second timeout limit on software HBA in ESXi. It appears whenever there is increased load on the cluster, like deep scrub or rebalance. Is it normal behaviour in production? Or is there something special we need to tune?
We are on latest Nautilus, 12 x 10 TB OSDs (4 servers), 25 Gbit/s Ethernet, erasure coded rbd pool with 128 PGs, aroun 200 PGs per OSD total.
ESXi Log:
2020-10-04T01:57:04.314Z cpu34:2098959)WARNING: iscsi_vmk: iscsivmk_ConnReceiveAtomic:517: vmhba64:CH:1 T:0 CN:0: Failed to receive data: Connection closed by peer 2020-10-04T01:57:04.314Z cpu34:2098959)iscsi_vmk: iscsivmk_ConnRxNotifyFailure:1235: vmhba64:CH:1 T:0 CN:0: Connection rx notifying failure: Failed to Receive. State=Bound 2020-10-04T01:57:04.566Z cpu19:2098979)WARNING: iscsi_vmk: iscsivmk_StopConnection:741: vmhba64:CH:1 T:0 CN:0: iSCSI connection is being marked "OFFLINE" (Event:4) 2020-10-04T01:57:04.654Z cpu7:2097866)WARNING: VMW_SATP_ALUA: satp_alua_issueCommandOnPath:788: Probe cmd 0xa3 failed for path "vmhba64:C2:T0:L0" (0x5/0x20/0x0). Check if failover mode is still ALUA.
OSD Log:
[303088.450088] Did not receive response to NOPIN on CID: 0, failing connection for I_T Nexus iqn.1994-05.com.redhat:esxi1,i,0x00023d000002,iqn.2003-01.com.redhat.iscsi-gw:iscsi-igw,t,0x01 [324926.694077] Did not receive response to NOPIN on CID: 0, failing connection for I_T Nexus iqn.1994-05.com.redhat:esxi2,i,0x00023d000001,iqn.2003-01.com.redhat.iscsi-gw:iscsi-igw,t,0x01 [407067.404538] ABORT_TASK: Found referenced iSCSI task_tag: 5891 [407076.077175] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 5891 [411677.887690] ABORT_TASK: Found referenced iSCSI task_tag: 6722 [411683.297425] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 6722 [481459.755876] ABORT_TASK: Found referenced iSCSI task_tag: 7930 [481460.787968] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 7930
Cheers, Martin_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Thanks! Does that mean that occasional iSCSI path drop-outs are somewhat expected? We are using SSDs for WAL/DB on each OSD server, so at least that. Do you think that If we buy additional 6/12 HDDs would that help with the IOPS for the VMs? Regards, Martin
On 4 Oct 2020, at 15:17, Martin Verges <martin.verges@croit.io> wrote:
Hello,
no iSCSI + VMware works without such problems.
We are on latest Nautilus, 12 x 10 TB OSDs (4 servers), 25 Gbit/s Ethernet, erasure coded rbd pool with 128 PGs, aroun 200 PGs per OSD total.
Nautilus is a good choice 12*10TB HDD is not good for VMs 25Gbit/s on HDD is way to much for that system 200 PGs per OSD is to much, I would suggest 75-100 PGs per OSD
You can improve latency on HDD clusters using external DB/WAL on NVMe. That might help you
-- Martin Verges Managing director
Mobile: +49 174 9335695 E-Mail: martin.verges@croit.io <mailto:martin.verges@croit.io> Chat: https://t.me/MartinVerges <https://t.me/MartinVerges>
croit GmbH, Freseniusstr. 31h, 81247 Munich CEO: Martin Verges - VAT-ID: DE310638492 Com. register: Amtsgericht Munich HRB 231263
Web: https://croit.io <https://croit.io/> YouTube: https://goo.gl/PGE1Bx <https://goo.gl/PGE1Bx>
Am So., 4. Okt. 2020 um 14:37 Uhr schrieb Golasowski Martin <martin.golasowski@vsb.cz <mailto:martin.golasowski@vsb.cz>>: Hi, does anyone here use CEPH iSCSI with VMware ESXi? It seems that we are hitting the 5 second timeout limit on software HBA in ESXi. It appears whenever there is increased load on the cluster, like deep scrub or rebalance. Is it normal behaviour in production? Or is there something special we need to tune?
We are on latest Nautilus, 12 x 10 TB OSDs (4 servers), 25 Gbit/s Ethernet, erasure coded rbd pool with 128 PGs, aroun 200 PGs per OSD total.
ESXi Log:
2020-10-04T01:57:04.314Z cpu34:2098959)WARNING: iscsi_vmk: iscsivmk_ConnReceiveAtomic:517: vmhba64:CH:1 T:0 CN:0: Failed to receive data: Connection closed by peer 2020-10-04T01:57:04.314Z cpu34:2098959)iscsi_vmk: iscsivmk_ConnRxNotifyFailure:1235: vmhba64:CH:1 T:0 CN:0: Connection rx notifying failure: Failed to Receive. State=Bound 2020-10-04T01:57:04.566Z cpu19:2098979)WARNING: iscsi_vmk: iscsivmk_StopConnection:741: vmhba64:CH:1 T:0 CN:0: iSCSI connection is being marked "OFFLINE" (Event:4) 2020-10-04T01:57:04.654Z cpu7:2097866)WARNING: VMW_SATP_ALUA: satp_alua_issueCommandOnPath:788: Probe cmd 0xa3 failed for path "vmhba64:C2:T0:L0" (0x5/0x20/0x0). Check if failover mode is still ALUA.
OSD Log:
[303088.450088] Did not receive response to NOPIN on CID: 0, failing connection for I_T Nexus iqn.1994-05.com.redhat:esxi1,i,0x00023d000002,iqn.2003-01.com.redhat.iscsi-gw:iscsi-igw,t,0x01 [324926.694077] Did not receive response to NOPIN on CID: 0, failing connection for I_T Nexus iqn.1994-05.com.redhat:esxi2,i,0x00023d000001,iqn.2003-01.com.redhat.iscsi-gw:iscsi-igw,t,0x01 [407067.404538] ABORT_TASK: Found referenced iSCSI task_tag: 5891 [407076.077175] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 5891 [411677.887690] ABORT_TASK: Found referenced iSCSI task_tag: 6722 [411683.297425] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 6722 [481459.755876] ABORT_TASK: Found referenced iSCSI task_tag: 7930 [481460.787968] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 7930
Cheers, Martin_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
Hello, in my personal opinion, HDDs are a technology from the last century and I would never ever think about using such old technology for modern VM/Container/... workloads. My time, as well as any employee is too precious to wait for a harddrive to find the requested data! Use EC on NVMe if you need to save some money. It's still much faster with lower latency than HDDs. As each HDD only adds like 100 IO/s and 20-30 MB/s to your cluster, you can throw in 100 Disks and won't even come near the performance of a single SSD. Yes, each disk will improve your performance, but by such a small amount that it makes no sense in my eyes.
Does that mean that occasional iSCSI path drop-outs are somewhat expected? Not that I'm aware of, but I have no HDD based ISCSI cluster at hand to check. Sorry.
-- Martin Verges Managing director Mobile: +49 174 9335695 E-Mail: martin.verges@croit.io Chat: https://t.me/MartinVerges croit GmbH, Freseniusstr. 31h, 81247 Munich CEO: Martin Verges - VAT-ID: DE310638492 Com. register: Amtsgericht Munich HRB 231263 Web: https://croit.io YouTube: https://goo.gl/PGE1Bx Am So., 4. Okt. 2020 um 16:06 Uhr schrieb Golasowski Martin < martin.golasowski@vsb.cz>:
Thanks!
Does that mean that occasional iSCSI path drop-outs are somewhat expected? We are using SSDs for WAL/DB on each OSD server, so at least that.
Do you think that If we buy additional 6/12 HDDs would that help with the IOPS for the VMs?
Regards, Martin
On 4 Oct 2020, at 15:17, Martin Verges <martin.verges@croit.io> wrote:
Hello,
no iSCSI + VMware works without such problems.
We are on latest Nautilus, 12 x 10 TB OSDs (4 servers), 25 Gbit/s Ethernet, erasure coded rbd pool with 128 PGs, aroun 200 PGs per OSD total.
Nautilus is a good choice 12*10TB HDD is not good for VMs 25Gbit/s on HDD is way to much for that system 200 PGs per OSD is to much, I would suggest 75-100 PGs per OSD
You can improve latency on HDD clusters using external DB/WAL on NVMe. That might help you
-- Martin Verges Managing director
Mobile: +49 174 9335695 E-Mail: martin.verges@croit.io Chat: https://t.me/MartinVerges
croit GmbH, Freseniusstr. 31h, 81247 Munich CEO: Martin Verges - VAT-ID: DE310638492 Com. register: Amtsgericht Munich HRB 231263
Web: https://croit.io YouTube: https://goo.gl/PGE1Bx
Am So., 4. Okt. 2020 um 14:37 Uhr schrieb Golasowski Martin < martin.golasowski@vsb.cz>:
Hi, does anyone here use CEPH iSCSI with VMware ESXi? It seems that we are hitting the 5 second timeout limit on software HBA in ESXi. It appears whenever there is increased load on the cluster, like deep scrub or rebalance. Is it normal behaviour in production? Or is there something special we need to tune?
We are on latest Nautilus, 12 x 10 TB OSDs (4 servers), 25 Gbit/s Ethernet, erasure coded rbd pool with 128 PGs, aroun 200 PGs per OSD total.
ESXi Log:
2020-10-04T01:57:04.314Z cpu34:2098959)WARNING: iscsi_vmk: iscsivmk_ConnReceiveAtomic:517: vmhba64:CH:1 T:0 CN:0: Failed to receive data: Connection closed by peer 2020-10-04T01:57:04.314Z cpu34:2098959)iscsi_vmk: iscsivmk_ConnRxNotifyFailure:1235: vmhba64:CH:1 T:0 CN:0: Connection rx notifying failure: Failed to Receive. State=Bound 2020-10-04T01:57:04.566Z cpu19:2098979)WARNING: iscsi_vmk: iscsivmk_StopConnection:741: vmhba64:CH:1 T:0 CN:0: iSCSI connection is being marked "OFFLINE" (Event:4) 2020-10-04T01:57:04.654Z cpu7:2097866)WARNING: VMW_SATP_ALUA: satp_alua_issueCommandOnPath:788: Probe cmd 0xa3 failed for path "vmhba64:C2:T0:L0" (0x5/0x20/0x0). Check if failover mode is still ALUA.
OSD Log:
[303088.450088] Did not receive response to NOPIN on CID: 0, failing connection for I_T Nexus iqn.1994-05.com.redhat:esxi1,i,0x00023d000002,iqn.2003-01.com.redhat.iscsi-gw:iscsi-igw,t,0x01 [324926.694077] Did not receive response to NOPIN on CID: 0, failing connection for I_T Nexus iqn.1994-05.com.redhat:esxi2,i,0x00023d000001,iqn.2003-01.com.redhat.iscsi-gw:iscsi-igw,t,0x01 [407067.404538] ABORT_TASK: Found referenced iSCSI task_tag: 5891 [407076.077175] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 5891 [411677.887690] ABORT_TASK: Found referenced iSCSI task_tag: 6722 [411683.297425] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 6722 [481459.755876] ABORT_TASK: Found referenced iSCSI task_tag: 7930 [481460.787968] ABORT_TASK: Sending TMR_FUNCTION_COMPLETE for ref_tag: 7930
Cheers, Martin_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On Sun, 4 Oct 2020, Martin Verges wrote:
Does that mean that occasional iSCSI path drop-outs are somewhat expected? Not that I'm aware of, but I have no HDD based ISCSI cluster at hand to check. Sorry.
I use iscsi extensively, but for ZFS and not ceph. Path drop-outs are not common; indeed, so far as I am aware, I have never had one. CentOS 7.8. Steve -- ---------------------------------------------------------------------------- Steve Thompson E-mail: smt AT vgersoft DOT com Voyager Software LLC Web: http://www DOT vgersoft DOT com 3901 N Charles St VSW Support: support AT vgersoft DOT com Baltimore MD 21218 "186,282 miles per second: it's not just a good idea, it's the law" ----------------------------------------------------------------------------
For clarity, the issue has been reported also before: https://www.spinics.net/lists/ceph-users/msg59798.html <https://www.spinics.net/lists/ceph-users/msg59798.html> https://www.spinics.net/lists/target-devel/msg10469.html
On 4 Oct 2020, at 16:46, Steve Thompson <smt@vgersoft.com> wrote:
On Sun, 4 Oct 2020, Martin Verges wrote:
Does that mean that occasional iSCSI path drop-outs are somewhat expected? Not that I'm aware of, but I have no HDD based ISCSI cluster at hand to check. Sorry.
I use iscsi extensively, but for ZFS and not ceph. Path drop-outs are not common; indeed, so far as I am aware, I have never had one. CentOS 7.8.
Steve -- ---------------------------------------------------------------------------- Steve Thompson E-mail: smt AT vgersoft DOT com Voyager Software LLC Web: http://www DOT vgersoft DOT com 3901 N Charles St VSW Support: support AT vgersoft DOT com Baltimore MD 21218 "186,282 miles per second: it's not just a good idea, it's the law" ----------------------------------------------------------------------------
Yep, and we're still experiencing it every few months. One (and only one) of our ESXi nodes, which are otherwise identical, is experiencing total freeze of all I/O, and it won't recover - I mean, ESXi is so dead, we have to go into IPMI and reset the box... We're using Croit's software, but the issue doesn't seem to be with CEPH so much as with vmware. That said, there's a couple of things you should be looking at: 1. Make sure you remember to set the RecoveryTimeout to 25 ? https://docs.ceph.com/en/latest/rbd/iscsi-initiator-esx/ 2. Make sure you have got working multipath across more than 1 adapter. What's possibly biting us right now, is that with 2 iscsi gateways in our cluster, and although both are autodiscovered at iscsi configuration time, we see that the ESXi nodes still only will show one path to each LUN. Currently these ESXi nodes have only 1 x 10gbit connected, it looks like I'll need to wire up the second connector and set up a second path to the iscsi gateway from that. It may not solve the problem, but it might lower the I/O on a single gateway enough that we won't see the problem anymore (and hopefully our customers stop getting pissed off). Cheers, Phil Golasowski Martin (martin.golasowski) writes:
For clarity, the issue has been reported also before:
https://www.spinics.net/lists/ceph-users/msg59798.html <https://www.spinics.net/lists/ceph-users/msg59798.html>
https://www.spinics.net/lists/target-devel/msg10469.html
On 4 Oct 2020, at 16:46, Steve Thompson <smt@vgersoft.com> wrote:
On Sun, 4 Oct 2020, Martin Verges wrote:
Does that mean that occasional iSCSI path drop-outs are somewhat expected? Not that I'm aware of, but I have no HDD based ISCSI cluster at hand to check. Sorry.
I use iscsi extensively, but for ZFS and not ceph. Path drop-outs are not common; indeed, so far as I am aware, I have never had one. CentOS 7.8.
Steve -- ---------------------------------------------------------------------------- Steve Thompson E-mail: smt AT vgersoft DOT com Voyager Software LLC Web: http://www DOT vgersoft DOT com 3901 N Charles St VSW Support: support AT vgersoft DOT com Baltimore MD 21218 "186,282 miles per second: it's not just a good idea, it's the law" ----------------------------------------------------------------------------
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
--
Oh, thanks, that does not sound very encouraging. In our case it looked the same, we had to reboot three ESXi nodes via IPMI, because it got stuck at ordinary soft reboot. 1. RecoveryTimeout is set at 25 on our nodes 2. We have one two-port adapter per node (Connect-X 5) and 4 iSCSI GWs total, one per OSD server. Multipath works, tested it by randomly rebooting one of two switches or manually shutting down the port. One curious thing I did not mention before is that we see a number of dropped Rx packets on each NIC that corresponds to the iSCSI VLAN. Increase in the dropped packets seems to correlate with the current IOPS load. I am beginning to settle with the version that our cluster is quite low on IOPS generally and slight increase in traffic may significantly raise latency on the iSCSI target. And ESXi is just being very touchy about that.
On 4 Oct 2020, at 18:59, Phil Regnauld <pr@x0.dk> wrote:
Yep, and we're still experiencing it every few months. One (and only one) of our ESXi nodes, which are otherwise identical, is experiencing total freeze of all I/O, and it won't recover - I mean, ESXi is so dead, we have to go into IPMI and reset the box...
We're using Croit's software, but the issue doesn't seem to be with CEPH so much as with vmware.
That said, there's a couple of things you should be looking at:
1. Make sure you remember to set the RecoveryTimeout to 25 ?
https://docs.ceph.com/en/latest/rbd/iscsi-initiator-esx/ <https://docs.ceph.com/en/latest/rbd/iscsi-initiator-esx/>
2. Make sure you have got working multipath across more than 1 adapter.
What's possibly biting us right now, is that with 2 iscsi gateways in our cluster, and although both are autodiscovered at iscsi configuration time, we see that the ESXi nodes still only will show one path to each LUN.
Currently these ESXi nodes have only 1 x 10gbit connected, it looks like I'll need to wire up the second connector and set up a second path to the iscsi gateway from that. It may not solve the problem, but it might lower the I/O on a single gateway enough that we won't see the problem anymore (and hopefully our customers stop getting pissed off).
Cheers, Phil
Golasowski Martin (martin.golasowski) writes:
For clarity, the issue has been reported also before:
https://www.spinics.net/lists/ceph-users/msg59798.html <https://www.spinics.net/lists/ceph-users/msg59798.html> <https://www.spinics.net/lists/ceph-users/msg59798.html <https://www.spinics.net/lists/ceph-users/msg59798.html>>
https://www.spinics.net/lists/target-devel/msg10469.html
On 4 Oct 2020, at 16:46, Steve Thompson <smt@vgersoft.com> wrote:
On Sun, 4 Oct 2020, Martin Verges wrote:
Does that mean that occasional iSCSI path drop-outs are somewhat expected? Not that I'm aware of, but I have no HDD based ISCSI cluster at hand to check. Sorry.
I use iscsi extensively, but for ZFS and not ceph. Path drop-outs are not common; indeed, so far as I am aware, I have never had one. CentOS 7.8.
Steve -- ---------------------------------------------------------------------------- Steve Thompson E-mail: smt AT vgersoft DOT com Voyager Software LLC Web: http://www DOT vgersoft DOT com 3901 N Charles St VSW Support: support AT vgersoft DOT com Baltimore MD 21218 "186,282 miles per second: it's not just a good idea, it's the law" ----------------------------------------------------------------------------
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
--
It is a load issue. Your combined load: client io, recovery, scrub is higher that what your cluster can handle. Whereas some ceph commands can block when things are very busy, VMWare iSCSI is less tolerant but it is not the problem. If you have charts, look at the metric for disk % utilization/busy and cpu % utilization, most probably your hdds were near 100% at the time of the issue, otherwise you can measure them manually at peak load. For the large majority of vm workloads, you really you should be using SSDs. If you have to use hdds, it could help to reduce the recovery and scrub load settings since you mentioned in your case the issue correlates to them, maybe try: osd_max_scrubs 1 osd_scrub_load_threshold 0.4 osd_scrub_sleep 0.2 osd_max_backfills 1 osd_recovery_max_active 1 osd_recovery_sleep 0.2 Generally it is beneficial to know what your peak workload is (iops/throughput) and what your cluster is capable of giving including during recovery and scrubbing. /Maged On 04/10/2020 19:47, Golasowski Martin wrote:
Oh, thanks, that does not sound very encouraging. In our case it looked the same, we had to reboot three ESXi nodes via IPMI, because it got stuck at ordinary soft reboot.
1. RecoveryTimeout is set at 25 on our nodes 2. We have one two-port adapter per node (Connect-X 5) and 4 iSCSI GWs total, one per OSD server. Multipath works, tested it by randomly rebooting one of two switches or manually shutting down the port. One curious thing I did not mention before is that we see a number of dropped Rx packets on each NIC that corresponds to the iSCSI VLAN. Increase in the dropped packets seems to correlate with the current IOPS load.
I am beginning to settle with the version that our cluster is quite low on IOPS generally and slight increase in traffic may significantly raise latency on the iSCSI target. And ESXi is just being very touchy about that.
On 4 Oct 2020, at 18:59, Phil Regnauld <pr@x0.dk <mailto:pr@x0.dk>> wrote:
Yep, and we're still experiencing it every few months. One (and only one) of our ESXi nodes, which are otherwise identical, is experiencing total freeze of all I/O, and it won't recover - I mean, ESXi is so dead, we have to go into IPMI and reset the box...
We're using Croit's software, but the issue doesn't seem to be with CEPH so much as with vmware.
That said, there's a couple of things you should be looking at:
1. Make sure you remember to set the RecoveryTimeout to 25 ?
https://docs.ceph.com/en/latest/rbd/iscsi-initiator-esx/
2. Make sure you have got working multipath across more than 1 adapter.
What's possibly biting us right now, is that with 2 iscsi gateways in our cluster, and although both are autodiscovered at iscsi configuration time, we see that the ESXi nodes still only will show one path to each LUN.
Currently these ESXi nodes have only 1 x 10gbit connected, it looks like I'll need to wire up the second connector and set up a second path to the iscsi gateway from that. It may not solve the problem, but it might lower the I/O on a single gateway enough that we won't see the problem anymore (and hopefully our customers stop getting pissed off).
Cheers, Phil
Golasowski Martin (martin.golasowski) writes:
For clarity, the issue has been reported also before:
https://www.spinics.net/lists/target-devel/msg10469.html
On 4 Oct 2020, at 16:46, Steve Thompson <smt@vgersoft.com> wrote:
On Sun, 4 Oct 2020, Martin Verges wrote:
Does that mean that occasional iSCSI path drop-outs are somewhat expected? Not that I'm aware of, but I have no HDD based ISCSI cluster at hand to check. Sorry.
I use iscsi extensively, but for ZFS and not ceph. Path drop-outs are not common; indeed, so far as I am aware, I have never had one. CentOS 7.8.
Steve -- ---------------------------------------------------------------------------- Steve Thompson E-mail: smt AT vgersoft DOT com Voyager Software LLC Web: http://www DOT vgersoft DOT com 3901 N Charles St VSW Support: support AT vgersoft DOT com Baltimore MD 21218 "186,282 miles per second: it's not just a good idea, it's the law" ----------------------------------------------------------------------------
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
--
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (5)
-
Golasowski Martin
-
Maged Mokhtar
-
Martin Verges
-
Phil Regnauld
-
Steve Thompson