All; We are in the middle of upgrading our primary cluster from 14.2.5 to 14.2.8. Our cluster utilizes 6 MDSs for 3 CephFS file systems. 3 MDSs are collocated with MON/MGR, and 3 MDSs are collocated with OSDs. At this point we have upgraded all 3 of the MON/MDS/MGR servers. The MDS on 2 of the 3 is currently not working, and we are seeing the below log messages. 2020-03-06 11:12:56.184 <> -1 mds.<daemon> unable to obtain rotating service keys; retrying 2020-03-06 11:13:26.184 <> 0 monclient: wait_auth_rotating timed out after 30 2020-03-06 11:13:26.184 <> -1 mds.<daemon> ERROR: failed to refresh rotating keys, maximum retry time reached. 2020-03-06 11:13:26.184 <> 1 mds.<daemon> suicide! Wanted state up:boot Any ideas? Thank you, Dominic L. Hilsbos, MBA Director - Information Technology Perform Air International Inc. DHilsbos@PerformAir.com www.PerformAir.com
On 3/6/20 7:46 PM, DHilsbos@performair.com wrote:
All;
We are in the middle of upgrading our primary cluster from 14.2.5 to 14.2.8. Our cluster utilizes 6 MDSs for 3 CephFS file systems. 3 MDSs are collocated with MON/MGR, and 3 MDSs are collocated with OSDs.
At this point we have upgraded all 3 of the MON/MDS/MGR servers. The MDS on 2 of the 3 is currently not working, and we are seeing the below log messages.
2020-03-06 11:12:56.184 <> -1 mds.<daemon> unable to obtain rotating service keys; retrying 2020-03-06 11:13:26.184 <> 0 monclient: wait_auth_rotating timed out after 30 2020-03-06 11:13:26.184 <> -1 mds.<daemon> ERROR: failed to refresh rotating keys, maximum retry time reached. 2020-03-06 11:13:26.184 <> 1 mds.<daemon> suicide! Wanted state up:boot
Any ideas?
Double check: Is the time correct on all the machines? cephx can have issues if there is a clock issue. Wido
Thank you,
Dominic L. Hilsbos, MBA Director - Information Technology Perform Air International Inc. DHilsbos@PerformAir.com www.PerformAir.com
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
All; When I went to check Wido's suggestion, I found the MDS daemons would start successfully. I obviously found no significant time differences. Sorry for making a mountain out of a mole-hill. Thank you, Dominic L. Hilsbos, MBA Director - Information Technology Perform Air International Inc. DHilsbos@PerformAir.com www.PerformAir.com -----Original Message----- From: Wido den Hollander [mailto:wido@42on.com] Sent: Friday, March 06, 2020 12:15 PM To: Dominic Hilsbos; ceph-users@ceph.io Subject: [ceph-users] Re: MDS Issues On 3/6/20 7:46 PM, DHilsbos@performair.com wrote:
All;
We are in the middle of upgrading our primary cluster from 14.2.5 to 14.2.8. Our cluster utilizes 6 MDSs for 3 CephFS file systems. 3 MDSs are collocated with MON/MGR, and 3 MDSs are collocated with OSDs.
At this point we have upgraded all 3 of the MON/MDS/MGR servers. The MDS on 2 of the 3 is currently not working, and we are seeing the below log messages.
2020-03-06 11:12:56.184 <> -1 mds.<daemon> unable to obtain rotating service keys; retrying 2020-03-06 11:13:26.184 <> 0 monclient: wait_auth_rotating timed out after 30 2020-03-06 11:13:26.184 <> -1 mds.<daemon> ERROR: failed to refresh rotating keys, maximum retry time reached. 2020-03-06 11:13:26.184 <> 1 mds.<daemon> suicide! Wanted state up:boot
Any ideas?
Double check: Is the time correct on all the machines? cephx can have issues if there is a clock issue. Wido
Thank you,
Dominic L. Hilsbos, MBA Director - Information Technology Perform Air International Inc. DHilsbos@PerformAir.com www.PerformAir.com
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (2)
-
DHilsbos@performair.com
-
Wido den Hollander