RGWs offline after upgrade to Nautilus
Hello, We have an RGW cluster that was recently upgraded from 12.2.11 to 14.2.22. The upgrade went mostly fine, though now several of our RGWs will not start. One RGW is working fine, the rest will not initialize. They are on a crash loop. This is part of a multisite configuration, and is currently not the master zone. Current master zone is running 14.2.22. These are the only two zones in the zonegroup. After turning debug up to 20, these are the log snippets between each crash: ``` 2023-07-20 14:29:56.371 7fd8dec40900 20 RGWRados::pool_iterate: got periods.1b6e1a93-98ba-4378-bc5c-d36cd5542f11.52 2023-07-20 14:29:56.371 7fd8dec40900 20 RGWRados::pool_iterate: got periods.1b6e1a93-98ba-4378-bc5c-d36cd5542f11.54 2023-07-20 14:29:56.371 7fd8dec40900 20 RGWRados::pool_iterate: got realms_names. <redacted> 2023-07-20 14:29:56.371 7fd8dec40900 20 RGWRados::pool_iterate: got <redacted> 2023-07-20 14:29:56.371 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.371 7fd8dec40900 20 rados_obj.operate() r=-2 bl.length=0 2023-07-20 14:29:56.371 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=-2 bl.length=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=-2 bl.length=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=46 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=114 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=46 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.374 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=686 2023-07-20 14:29:56.374 7fd8dec40900 20 period zonegroup init ret 0 2023-07-20 14:29:56.374 7fd8dec40900 20 period zonegroup name <redacted> 2023-07-20 14:29:56.374 7fd8dec40900 20 using current period zonegroup <redacted> 2023-07-20 14:29:56.374 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.374 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=46 2023-07-20 14:29:56.374 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.375 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=903 2023-07-20 14:29:56.375 7fd8dec40900 10 Cannot find current period zone using local zone 2023-07-20 14:29:56.375 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.375 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=903 2023-07-20 14:29:56.375 7fd8dec40900 20 zone <redacted> 2023-07-20 14:29:56.375 7fd8dec40900 20 generating connection object for zone <redacted> id f10b465f-bf18-47d0-a51c-ca4f17118ee1 2023-07-20 14:34:56.198 7fd8cafe8700 -1 Initialization timeout, failed to initialize ``` I’ve checked all file permissions, filesystem free space, disabled selinux and firewalld, tried turning up the initialization timeout to 600, and tried removing all non-essential config from ceph.conf. All produce the same results. I would greatly appreciate any other ideas or insight. Thanks, Ben
Hi, a couple of threads with similar error messages all lead back to some sort of pool or osd issue. What is your current cluster status (ceph -s)? Do you have some full OSDs? Those can cause this initialization timeout as well as hit the max_pg_per_osd limit. So a few more cluster details could help here. Thanks, Eugen Zitat von "Ben.Zieglmeier" <Ben.Zieglmeier@target.com>:
Hello,
We have an RGW cluster that was recently upgraded from 12.2.11 to 14.2.22. The upgrade went mostly fine, though now several of our RGWs will not start. One RGW is working fine, the rest will not initialize. They are on a crash loop. This is part of a multisite configuration, and is currently not the master zone. Current master zone is running 14.2.22. These are the only two zones in the zonegroup. After turning debug up to 20, these are the log snippets between each crash: ``` 2023-07-20 14:29:56.371 7fd8dec40900 20 RGWRados::pool_iterate: got periods.1b6e1a93-98ba-4378-bc5c-d36cd5542f11.52 2023-07-20 14:29:56.371 7fd8dec40900 20 RGWRados::pool_iterate: got periods.1b6e1a93-98ba-4378-bc5c-d36cd5542f11.54 2023-07-20 14:29:56.371 7fd8dec40900 20 RGWRados::pool_iterate: got realms_names. <redacted> 2023-07-20 14:29:56.371 7fd8dec40900 20 RGWRados::pool_iterate: got <redacted> 2023-07-20 14:29:56.371 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.371 7fd8dec40900 20 rados_obj.operate() r=-2 bl.length=0 2023-07-20 14:29:56.371 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=-2 bl.length=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=-2 bl.length=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=46 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=114 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.373 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=46 2023-07-20 14:29:56.373 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.374 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=686 2023-07-20 14:29:56.374 7fd8dec40900 20 period zonegroup init ret 0 2023-07-20 14:29:56.374 7fd8dec40900 20 period zonegroup name <redacted> 2023-07-20 14:29:56.374 7fd8dec40900 20 using current period zonegroup <redacted> 2023-07-20 14:29:56.374 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.374 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=46 2023-07-20 14:29:56.374 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.375 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=903 2023-07-20 14:29:56.375 7fd8dec40900 10 Cannot find current period zone using local zone 2023-07-20 14:29:56.375 7fd8dec40900 20 rados->read ofs=0 len=0 2023-07-20 14:29:56.375 7fd8dec40900 20 rados_obj.operate() r=0 bl.length=903 2023-07-20 14:29:56.375 7fd8dec40900 20 zone <redacted> 2023-07-20 14:29:56.375 7fd8dec40900 20 generating connection object for zone <redacted> id f10b465f-bf18-47d0-a51c-ca4f17118ee1 2023-07-20 14:34:56.198 7fd8cafe8700 -1 Initialization timeout, failed to initialize ```
I’ve checked all file permissions, filesystem free space, disabled selinux and firewalld, tried turning up the initialization timeout to 600, and tried removing all non-essential config from ceph.conf. All produce the same results. I would greatly appreciate any other ideas or insight.
Thanks, Ben _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Changing sending email address as something was wrong with my last one. Still OP here. Cluster is generally healthy. Not running out of storage space or pools filling up. As mentioned in the original post, one RGW is able to come online. I've cross-compared about every file permission, config file, keyring, etc between the working RGW and all other non working RGWs, and nothing seems to allow them to rejoin the cluster. ceph -s [root@host ceph]# ceph -s cluster: id: <id> health: HEALTH_WARN 601 large omap objects 502 pgs not deep-scrubbed in time 1 pgs not scrubbed in time services: mon: 3 daemons, quorum <mon1>,<mon2>,<mon3> (age 28h) mgr: <mgr1>(active, since 28h), standbys: <mgr2>, <mgr3> osd: 130 osds: 130 up (since 3d), 130 in rgw: 1 daemon active (<rgw1>) task status: data: pools: 7 pools, 4288 pgs objects: 926.15M objects, 88 TiB usage: 397 TiB used, 646 TiB / 1.0 PiB avail pgs: 4258 active+clean 30 active+clean+scrubbing+deep io: client: 340 KiB/s rd, 280 KiB/s wr, 370 op/s rd, 496 op/s wr ceph df: RAW STORAGE: CLASS SIZE AVAIL USED RAW USED %RAW USED hdd 763 TiB 450 TiB 313 TiB 313 TiB 41.04 ssd 279 TiB 196 TiB 80 TiB 84 TiB 29.95 TOTAL 1.0 PiB 646 TiB 394 TiB 397 TiB 38.07 POOLS: POOL ID PGS STORED OBJECTS USED %USED MAX AVAIL .rgw.root 51 32 172 KiB 98 14 MiB 0 177 TiB zone.rgw.control 60 32 0 B 8 0 B 0 177 TiB zone.rgw.meta 61 32 11 MiB 34.04k 5.0 GiB 0 177 TiB zone.rgw.log 62 32 508 GiB 438.39k 508 GiB 0.09 177 TiB zone.rgw.buckets.data 63 4096 88 TiB 925.20M 361 TiB 40.47 177 TiB zone.rgw.buckets.index 64 32 890 GiB 469.31k 890 GiB 0.16 177 TiB zone.rgw.buckets.non-ec 66 32 3.7 MiB 610 3.7 MiB 0 177 TiB
Hi, apparently, my previous suggestions don't apply here (full OSDs or max_pgs_per_osd limit). Did you also check the rgw client keyrings? Did you also upgrade the operating system? Maybe some apparmor stuff? Can you set debug to 30 to see if there're more to see? Anything in the mon or mgr logs or in the syslog? Thanks, Eugen Zitat von bzieglmeier@gmail.com:
Changing sending email address as something was wrong with my last one. Still OP here.
Cluster is generally healthy. Not running out of storage space or pools filling up. As mentioned in the original post, one RGW is able to come online. I've cross-compared about every file permission, config file, keyring, etc between the working RGW and all other non working RGWs, and nothing seems to allow them to rejoin the cluster.
ceph -s [root@host ceph]# ceph -s cluster: id: <id> health: HEALTH_WARN 601 large omap objects 502 pgs not deep-scrubbed in time 1 pgs not scrubbed in time
services: mon: 3 daemons, quorum <mon1>,<mon2>,<mon3> (age 28h) mgr: <mgr1>(active, since 28h), standbys: <mgr2>, <mgr3> osd: 130 osds: 130 up (since 3d), 130 in rgw: 1 daemon active (<rgw1>)
task status:
data: pools: 7 pools, 4288 pgs objects: 926.15M objects, 88 TiB usage: 397 TiB used, 646 TiB / 1.0 PiB avail pgs: 4258 active+clean 30 active+clean+scrubbing+deep
io: client: 340 KiB/s rd, 280 KiB/s wr, 370 op/s rd, 496 op/s wr
ceph df: RAW STORAGE: CLASS SIZE AVAIL USED RAW USED %RAW USED hdd 763 TiB 450 TiB 313 TiB 313 TiB 41.04 ssd 279 TiB 196 TiB 80 TiB 84 TiB 29.95 TOTAL 1.0 PiB 646 TiB 394 TiB 397 TiB 38.07
POOLS: POOL ID PGS STORED OBJECTS USED %USED MAX AVAIL .rgw.root 51 32 172 KiB 98 14 MiB 0 177 TiB zone.rgw.control 60 32 0 B 8 0 B 0 177 TiB zone.rgw.meta 61 32 11 MiB 34.04k 5.0 GiB 0 177 TiB zone.rgw.log 62 32 508 GiB 438.39k 508 GiB 0.09 177 TiB zone.rgw.buckets.data 63 4096 88 TiB 925.20M 361 TiB 40.47 177 TiB zone.rgw.buckets.index 64 32 890 GiB 469.31k 890 GiB 0.16 177 TiB zone.rgw.buckets.non-ec 66 32 3.7 MiB 610 3.7 MiB 0 177 TiB _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (3)
-
Ben.Zieglmeier
-
bzieglmeier@gmail.com
-
Eugen Block