Module 'cephadm' has failed: invalid literal for int() with base 10:
Hi, Our Ceph cluster is in an error state with the message: # ceph status cluster: id: 58140ed2-4ed4-11ed-b4db-5c6f69756a60 health: HEALTH_ERR Module 'cephadm' has failed: invalid literal for int() with base 10: '352.broken' This happened after trying to re-add an OSD which had failed. Adopting it back in to the Ceph failed because a directory was causing problems in /var/lib/ceph/{cephid}/osd.352. To re-add the OSD I renamed it to osd.352.broken (rather than delete it), re-ran the command and then everything worked perfectly. Then 5 minutes later the ceph orchestrator went into "HEALTH_ERR" I've removed that directory, but "cephadm" isn't cleaning up after itself. Does anyone know if there's a way I can clear the cache for this directory it's tried to inventory and failed? Thanks, Duncan -- Dr Duncan Tooke | Research Cluster Administrator Centre for Computational Biology, Weatherall Institute of Molecular Medicine, University of Oxford, OX3 9DS www.imm.ox.ac.uk<http://www.imm.ox.ac.uk>
Hi, have you tried a mgr failover? Zitat von Duncan M Tooke <duncan.tooke@imm.ox.ac.uk>:
Hi,
Our Ceph cluster is in an error state with the message:
# ceph status cluster: id: 58140ed2-4ed4-11ed-b4db-5c6f69756a60 health: HEALTH_ERR Module 'cephadm' has failed: invalid literal for int() with base 10: '352.broken'
This happened after trying to re-add an OSD which had failed. Adopting it back in to the Ceph failed because a directory was causing problems in /var/lib/ceph/{cephid}/osd.352. To re-add the OSD I renamed it to osd.352.broken (rather than delete it), re-ran the command and then everything worked perfectly. Then 5 minutes later the ceph orchestrator went into "HEALTH_ERR"
I've removed that directory, but "cephadm" isn't cleaning up after itself. Does anyone know if there's a way I can clear the cache for this directory it's tried to inventory and failed?
Thanks,
Duncan -- Dr Duncan Tooke | Research Cluster Administrator Centre for Computational Biology, Weatherall Institute of Molecular Medicine, University of Oxford, OX3 9DS www.imm.ox.ac.uk<http://www.imm.ox.ac.uk>
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi, Sorry, I didn't update this discussion yesterday. That was indeed exactly what was required, and it immediately recovered. Good shout 😊 Best wishes, Duncan -- Dr Duncan Tooke | Research Cluster Administrator Centre for Computational Biology, Weatherall Institute of Molecular Medicine, University of Oxford, OX3 9DS www.imm.ox.ac.uk
-----Original Message----- From: Eugen Block <eblock@nde.ag> Sent: Wednesday, April 12, 2023 8:10 AM To: ceph-users@ceph.io Subject: [ceph-users] Re: Module 'cephadm' has failed: invalid literal for int() with base 10:
Hi,
have you tried a mgr failover?
Zitat von Duncan M Tooke <duncan.tooke@imm.ox.ac.uk>:
Hi,
Our Ceph cluster is in an error state with the message:
# ceph status cluster: id: 58140ed2-4ed4-11ed-b4db-5c6f69756a60 health: HEALTH_ERR Module 'cephadm' has failed: invalid literal for int() with base 10: '352.broken'
This happened after trying to re-add an OSD which had failed. Adopting it back in to the Ceph failed because a directory was causing problems in /var/lib/ceph/{cephid}/osd.352. To re-add the OSD I renamed it to osd.352.broken (rather than delete it), re-ran the command and then everything worked perfectly. Then 5 minutes later the ceph orchestrator went into "HEALTH_ERR"
I've removed that directory, but "cephadm" isn't cleaning up after itself. Does anyone know if there's a way I can clear the cache for this directory it's tried to inventory and failed?
Thanks,
Duncan -- Dr Duncan Tooke | Research Cluster Administrator Centre for Computational Biology, Weatherall Institute of Molecular Medicine, University of Oxford, OX3 9DS www.imm.ox.ac.uk<http://www.imm.ox.ac.uk>
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (2)
-
Duncan M Tooke
-
Eugen Block