Orch stops refreshing daemon list
I have a problem following a 20.2.3 -> 20.2.4 upgrade. This has happened once on a virtualised test cluster, and now on a small production cluster. After it happened on test, I rolled all nodes back to snapshots and tried again, and it worked, so I assumed finger trouble and proceeded to the production cluster. But now it is stuck... - The upgrade itself is completely successful. Repeated "ceph orch ps --refresh" looks good and shows the "refreshed" value returning to zero each time. - I then rotate the SSH keys, and everything at first appears ok. "check-host" works between every pair of hosts in each direction. - But... "ceph orch ps --refresh" has now stopped refreshing. And any attempts to restart or redeploy a daemon all say it is scheduled but nothing happens. I've tried failing MGRs, rebooting a node then failing back to the MGR on it. There's nothing odd that I can spot in the MGR logs. After the reboot, "ceph orch ps --refresh" still shows the OLD container IDs, and an ever-increasing last-refresh time. I've also rotated the client.admin key on the production cluster without any problems. I don't think that is causing this orch problem, as the test cluster originally failed before client.admin had been rotated. Can anyone suggest how I may troubleshoot this? Regards, Chris
participants (1)
-
Chris Palmer