Hello, we are configure new ceph cluster with Mellanox 2x100Gbps cards. We bond this two ports to MLAG bond0 interface. In the async+posix mode everythink is OK, cluster is in the HELTH_OK state. CEPH version is 18.2.1. Then we tried to configure RoCE for cluster part of network, but without success. Our ceph config dump (only relevant config): global advanced ms_async_rdma_device_name mlx5_bond_0 * global advanced ms_async_rdma_gid_idx 3 global host:ceph1-nvme advanced ms_async_rdma_local_gid 0000:0000:0000:0000:0000:ffff:a0d9:05d8 * global host:ceph2-nvme advanced ms_async_rdma_local_gid 0000:0000:0000:0000:0000:ffff:a0d9:05d7 * global host:ceph3-nvme advanced ms_async_rdma_local_gid 0000:0000:0000:0000:0000:ffff:a0d9:05d6 * global advanced ms_async_rdma_roce_ver 2 global advanced ms_async_rdma_type rdma * global advanced ms_cluster_type async+rdma * global advanced ms_public_type async+posix * On the ceph1-nvme there is this show_gids.sh list: # ./show_gids.sh DEV PORT INDEX GID IPv4 VER DEV --- ---- ----- --- ------------ --- --- mlx5_bond_0 1 0 fe80:0000:0000:0000:0e42:a1ff:fe93:b004 v1 bond0 mlx5_bond_0 1 1 fe80:0000:0000:0000:0e42:a1ff:fe93:b004 v2 bond0 mlx5_bond_0 1 2 0000:0000:0000:0000:0000:ffff:a0d9:05d8 160.217.5.216 v1 bond0 mlx5_bond_0 1 3 0000:0000:0000:0000:0000:ffff:a0d9:05d8 160.217.5.216 v2 bond0 n_gids_found=4 I have set this line in /etc/security/limits.conf: * hard memlock unlimited But when I tried to restart ceph.target, OSD nodes didn't start with this errors, see attachment. Mellanox drivers are from Debian bookworm kernel. Is there somethink missing in config, or some errors? When I change ms_cluster_type to async+posix and restart ceph.target, cluster converged to HEALTH_OK state... Thanks for advices... Sincerely Jan Marek -- Ing. Jan Marek University of South Bohemia Academic Computer Centre Phone: +420389032080 http://www.gnu.org/philosophy/no-word-attachments.cs.html