ceph 16.2.10 cluster down
Hello, I'm currently investigating a downed ceph cluster that I cannot communicate with. Setup is: 3 hosts with each 12 disks (osd/mon) 3 vm's with mon/mds/mgr The vm's are unavailable at the moment and one of the hosts is online with osd/mon running. When issuing the command ceph -s nothing happens and after 5 minutes the following. 2023-01-26T12:40:08.111+0100 7f68b4b8b700 0 monclient(hunting): authenticate timed out after 300 What would be the best way of troubleshooting this? Venlig hilsen - Mit freundlichen Grüßen - Kind Regards, Jens Galsgaard
Hi, On 26.01.23 12:46, Jens Galsgaard wrote:
Setup is: 3 hosts with each 12 disks (osd/mon) 3 vm's with mon/mds/mgr
The vm's are unavailable at the moment and one of the hosts is online with osd/mon running.
You have only one out of six MONs running. This MON is unable to form a quorum. Best would be to start the other MONs so that you have at least 4 running. They could form a quorum and then the cluster will respond again. If that is not possible and you want to recover from a catastrophic failure you need to manually edit the MON map (with monmaptool) and remove all but the running MON from it. Then this MON will only see itself as active in the cluster and form the quorum. https://docs.ceph.com/en/quincy/rados/troubleshooting/troubleshooting-mon https://docs.ceph.com/en/quincy/rados/operations/add-or-rm-mons/#removing-mo... Regards -- Robert Sander Heinlein Consulting GmbH Schwedter Str. 8/9b, 10119 Berlin https://www.heinlein-support.de Tel: 030 / 405051-43 Fax: 030 / 405051-19 Amtsgericht Berlin-Charlottenburg - HRB 220009 B Geschäftsführer: Peer Heinlein - Sitz: Berlin
Hi Robert, After removing the dead monitors with the monmaptool the mon container has vanished from podman. So this somehow made things worse. Is it possible to create and add new monitors? Re-bootstrap the cluster in lack of better terms? //Jens -----Oprindelig meddelelse----- Fra: Robert Sander <r.sander@heinlein-support.de> Sendt: Thursday, January 26, 2023 1:25 PM Til: ceph-users@ceph.io Emne: [ceph-users] Re: ceph 16.2.10 cluster down Hi, On 26.01.23 12:46, Jens Galsgaard wrote:
Setup is: 3 hosts with each 12 disks (osd/mon) 3 vm's with mon/mds/mgr
The vm's are unavailable at the moment and one of the hosts is online with osd/mon running.
You have only one out of six MONs running. This MON is unable to form a quorum. Best would be to start the other MONs so that you have at least 4 running. They could form a quorum and then the cluster will respond again. If that is not possible and you want to recover from a catastrophic failure you need to manually edit the MON map (with monmaptool) and remove all but the running MON from it. Then this MON will only see itself as active in the cluster and form the quorum. https://docs.ceph.com/en/quincy/rados/troubleshooting/troubleshooting-mon https://docs.ceph.com/en/quincy/rados/operations/add-or-rm-mons/#removing-mo... Regards -- Robert Sander Heinlein Consulting GmbH Schwedter Str. 8/9b, 10119 Berlin https://www.heinlein-support.de Tel: 030 / 405051-43 Fax: 030 / 405051-19 Amtsgericht Berlin-Charlottenburg - HRB 220009 B Geschäftsführer: Peer Heinlein - Sitz: Berlin _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Jens, On 26.01.23 16:17, Jens Galsgaard wrote:
After removing the dead monitors with the monmaptool the mon container has vanished from podman. So this somehow made things worse.
You have not mentioned that you are running Ceph in containers. The procedure to repair the MON map may look a little bit different then. The documentation only shows the non-containerized procedure. Do you still have the systemd unit for the last working MON? Run "systemctl --all | grep mon" to look for it. If it is stop try to start it. Does the MON run?
Is it possible to create and add new monitors? Re-bootstrap the cluster in lack of better terms?
It should be possible to restore the MON db using an exisiting OSD: https://docs.ceph.com/en/pacific/rados/troubleshooting/troubleshooting-mon/#... This procedure would also have to be adapted for a containerized setup. Regards -- Robert Sander Heinlein Consulting GmbH Schwedter Str. 8/9b, 10119 Berlin https://www.heinlein-support.de Tel: 030 / 405051-43 Fax: 030 / 405051-19 Amtsgericht Berlin-Charlottenburg - HRB 220009 B Geschäftsführer: Peer Heinlein - Sitz: Berlin
I got the monitor started with cephadm and could run ceph-mon manually to get the mon service running again. Thanks for the hints along the way Robert 😊 //Jens -----Oprindelig meddelelse----- Fra: Robert Sander <r.sander@heinlein-support.de> Sendt: Thursday, January 26, 2023 4:23 PM Til: ceph-users@ceph.io Emne: [ceph-users] Re: ceph 16.2.10 cluster down Hi Jens, On 26.01.23 16:17, Jens Galsgaard wrote:
After removing the dead monitors with the monmaptool the mon container has vanished from podman. So this somehow made things worse.
You have not mentioned that you are running Ceph in containers. The procedure to repair the MON map may look a little bit different then. The documentation only shows the non-containerized procedure. Do you still have the systemd unit for the last working MON? Run "systemctl --all | grep mon" to look for it. If it is stop try to start it. Does the MON run?
Is it possible to create and add new monitors? Re-bootstrap the cluster in lack of better terms?
It should be possible to restore the MON db using an exisiting OSD: https://docs.ceph.com/en/pacific/rados/troubleshooting/troubleshooting-mon/#... This procedure would also have to be adapted for a containerized setup. Regards -- Robert Sander Heinlein Consulting GmbH Schwedter Str. 8/9b, 10119 Berlin https://www.heinlein-support.de Tel: 030 / 405051-43 Fax: 030 / 405051-19 Amtsgericht Berlin-Charlottenburg - HRB 220009 B Geschäftsführer: Peer Heinlein - Sitz: Berlin _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (2)
-
Jens Galsgaard
-
Robert Sander