Hi Kirby, These failing OSDs have been a part of the cluster from the beginning. Another group of OSDs are added later to the cluster (before upgrade) and have no problem. The cluster is upgraded from Reef 18.2.4. I am sharing the recent osdmap and the older one (the one failing OSDs used) Current epoch 72555 : https://pastes.io/epoch-72555 Older epoch that down osd uses: https://pastes.io/epoch-70297 On the older epoch all OSDs are UP, but the recent epoch has 24 OSDs down. Thanks, Huseyin hcotuk@gmail.com
On 24 Oct 2025, at 18:03, Kirby Haze <kirbyhaze01@gmail.com> wrote:
Were these osds created before or after the upgrade? And from which reef version did you upgrade from?
Could you try to get the latest osdmap the down osd has? What does that look like?
On Fri, Oct 24, 2025 at 7:42 AM Huseyin Cotuk <hcotuk@gmail.com <mailto:hcotuk@gmail.com>> wrote:
More debug logs from a failing OSD:
Oct 24 17:39:52 ank-backup01 ceph-4e7e7d1c-22db-49c7-9f24-5a75cd3a3b9f-osd-11[4100566]: 2025-10-24T14:39:52.271+0000 77f47e35c640 20 osd.11 70297 reports for 0 queries Oct 24 17:39:52 ank-backup01 ceph-4e7e7d1c-22db-49c7-9f24-5a75cd3a3b9f-osd-11[4100566]: 2025-10-24T14:39:52.271+0000 77f47e35c640 15 osd.11 70297 collect_pg_stats Oct 24 17:39:52 ank-backup01 ceph-osd[4100574]: osd.11 pg_epoch: 70297 pg[23.b9s0( v 69083'8392506 (68272'8387884,69083'8392506] local-lis/les=69076/69077 n=8314181 ec=21807/21807 lis/c=69076/68833 les/c/f=69077/68834/39143 sis=70280) [11,NONE,NONE,12,31,21,7,NONE,NONE,26,0]p11(0) r=0 lpr=70297 pi=[68833,70280)/5 crt=69083'8392506 lcod 0'0 mlcod 0'0 unknown mbc={}] PeeringState::prepare_stats_for_publish reporting purged_snaps [] Oct 24 17:39:52 ank-backup01 ceph-4e7e7d1c-22db-49c7-9f24-5a75cd3a3b9f-osd-11[4100566]: 2025-10-24T14:39:52.271+0000 77f47e35c640 20 osd.11 pg_epoch: 70297 pg[23.b9s0( v 69083'8392506 (68272'8387884,69083'8392506] local-lis/les=69076/69077 n=8314181 ec=21807/21807 lis/c=69076/68833 les/c/f=69077/68834/39143 sis=70280) [11,NONE,NONE,12,31,21,7,NONE,NONE,26,0]p11(0) r=0 lpr=70297 pi=[68833,70280)/5 crt=69083'8392506 lcod 0'0 mlcod 0'0 unknown mbc={}] PeeringState::prepare_stats_for_publish reporting purged_snaps [] Oct 24 17:39:52 ank-backup01 ceph-osd[4100574]: osd.11 pg_epoch: 70297 pg[23.b9s0( v 69083'8392506 (68272'8387884,69083'8392506] local-lis/les=69076/69077 n=8314181 ec=21807/21807 lis/c=69076/68833 les/c/f=69077/68834/39143 sis=70280) [11,NONE,NONE,12,31,21,7,NONE,NONE,26,0]p11(0) r=0 lpr=70297 pi=[68833,70280)/5 crt=69083'8392506 lcod 0'0 mlcod 0'0 unknown mbc={}] PeeringState::prepare_stats_for_publish publish_stats_to_osd 70297:52872808 Oct 24 17:39:52 ank-backup01 ceph-4e7e7d1c-22db-49c7-9f24-5a75cd3a3b9f-osd-11[4100566]: 2025-10-24T14:39:52.271+0000 77f47e35c640 15 osd.11 pg_epoch: 70297 pg[23.b9s0( v 69083'8392506 (68272'8387884,69083'8392506] local-lis/les=69076/69077 n=8314181 ec=21807/21807 lis/c=69076/68833 les/c/f=69077/68834/39143 sis=70280) [11,NONE,NONE,12,31,21,7,NONE,NONE,26,0]p11(0) r=0 lpr=70297 pi=[68833,70280)/5 crt=69083'8392506 lcod 0'0 mlcod 0'0 unknown mbc={}] PeeringState::prepare_stats_for_publish publish_stats_to_osd 70297:52872808 Oct 24 17:39:52 ank-backup01 ceph-osd[4100574]: osd.11 pg_epoch: 70297 pg[23.5fs2( v 69081'8395094 (68276'8390628,69081'8395094] local-lis/les=69078/69079 n=8316950 ec=21807/21807 lis/c=69078/68833 les/c/f=69079/68834/39143 sis=70280) [NONE,NONE,11,8,14,16,0,20,17,18,NONE]p11(2) r=2 lpr=70297 pi=[68833,70280)/7 crt=69081'8395094 lcod 0'0 mlcod 0'0 unknown mbc={}] PeeringState::prepare_stats_for_publish reporting purged_snaps []
Selamlar, Huseyin Cotuk hcotuk@gmail.com <mailto:hcotuk@gmail.com>
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>