All,
The host, reesi002, in the LRC has what appears to be some sort of storage issue. All operations on the host, including basic filesystem operations, are extremely slow. Many of the OSDs have already crashed and been down for days, possibly
weeks.
I had intended on slowly removing the OSDs from the cluster in an attempt to let the iSCSI gateway do its thing with less load on the host but just purging one OSD from the cluster (osd.5) has caused the other OSDs on the host to bounce
up and down.
Since reesi002 is already in a state where it could crash at any moment (as has already happened this week), I am going to proceed with taking the remaining daemons on the host out of the cluster. There is allegedly a path that should
allow me to move the iSCSI LUNs serving the RHV VMs to another host non-disruptively but we are all familiar with our RHV experience thus far.
This may result in disruption to the lab including teuthology runs, DNS, ceph.io, lists.ceph.io, etc.
- David