Hi, for some reason radosgw stopped working. Cluster status: [root@ctplmon1 ~]# ceph -v ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable) [root@ctplmon1 ~]# ceph -s cluster: id: 0a6e5422-ac75-4093-af20-528ee00cc847 health: HEALTH_ERR 6 OSD(s) experiencing slow operations in BlueStore 2 backfillfull osd(s) 1 full osd(s) 1 nearfull osd(s) Low space hindering backfill (add storage if this doesn't resolve itself): 32 pgs backfill_toofull Degraded data redundancy: 835306/1285383707 objects degraded (0.065%), 6 pgs degraded, 5 pgs undersized 76 pgs not deep-scrubbed in time 45 pgs not scrubbed in time Full OSDs blocking recovery: 1 pg recovery_toofull 9 pool(s) full 9 daemons have recently crashed services: mon: 3 daemons, quorum ctplmon1,ctplmon3,ctplmon2 (age 36m) mgr: ctplmon1(active, since 65m) mds: 1/1 daemons up osd: 193 osds: 191 up (since 8m), 191 in (since 9m); 267 remapped pgs rgw: 2 daemons active (1 hosts, 1 zones) data: volumes: 1/1 healthy pools: 10 pools, 793 pgs objects: 257.08M objects, 292 TiB usage: 614 TiB used, 386 TiB / 1000 TiB avail pgs: 835306/1285383707 objects degraded (0.065%) 225512620/1285383707 objects misplaced (17.544%) 525 active+clean 230 active+remapped+backfilling 32 active+remapped+backfill_toofull 5 active+undersized+degraded+remapped+backfilling 1 active+recovery_toofull+degraded io: recovery: 978 MiB/s, 825 objects/s --- Do not know if it is related but the cluster has been rebalancing for a few days now, after I've set EC pool only to use hdd. --- If I start rgw with debug I get something like this in logs: [root@ctplmon2 ~]# radosgw -c /etc/ceph/ceph.conf --setuser ceph --setgroup ceph -n client.radosgw.moja.shramba.ctplmon2 -f -m 194.249.4.104:6789 --debug-rgw=99/99 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize In logs I get: 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 deferred set uid:gid to 167:167 (ceph:ceph) 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable), process radosgw, pid 168935 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework: beast 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework conf key: port, val: 4444 2024-12-21T23:16:59.898+0100 7f65a19257c0 1 radosgw_Main not setting numa affinity 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 rgw_d3n: rgw_d3n_l1_local_datacache_enabled=0 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 D3N datacache enabled: 0 2024-12-21T23:16:59.901+0100 7f658dffb640 20 reqs_thread_entry: start 2024-12-21T23:16:59.901+0100 7f658d7fa640 10 entry start 2024-12-21T23:16:59.908+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: realm 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 4 rgw main: RGWPeriod::init failed to init realm id : (2) No such file or directory 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: searching for the correct realm 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.215+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 zone default found 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Realm: () 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: ZoneGroup: default (b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578) 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Zone: default (c2c70444-7a41-4acd-a0d0-9f87d324ec72) 2024-12-21T23:17:00.295+0100 7f65a19257c0 10 cannot find current period zonegroup using local zonegroup configuration 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: zonegroup default 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started sync module instance, tier type = 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started zone id=c2c70444-7a41-4acd-a0d0-9f87d324ec72 (name=default) with tier type = 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize --- Any ideas what might cause rgw to stop working? Kind regards, Rok
The full OSD is most likely the reason. You can temporarily increase the threshold to 0.97 or so, but you need to prevent that to happen. The cluster usually starts warning you at 85%. Zitat von Rok Jaklič <rjaklic@gmail.com>:
Hi,
for some reason radosgw stopped working.
Cluster status: [root@ctplmon1 ~]# ceph -v ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable) [root@ctplmon1 ~]# ceph -s cluster: id: 0a6e5422-ac75-4093-af20-528ee00cc847 health: HEALTH_ERR 6 OSD(s) experiencing slow operations in BlueStore 2 backfillfull osd(s) 1 full osd(s) 1 nearfull osd(s) Low space hindering backfill (add storage if this doesn't resolve itself): 32 pgs backfill_toofull Degraded data redundancy: 835306/1285383707 objects degraded (0.065%), 6 pgs degraded, 5 pgs undersized 76 pgs not deep-scrubbed in time 45 pgs not scrubbed in time Full OSDs blocking recovery: 1 pg recovery_toofull 9 pool(s) full 9 daemons have recently crashed
services: mon: 3 daemons, quorum ctplmon1,ctplmon3,ctplmon2 (age 36m) mgr: ctplmon1(active, since 65m) mds: 1/1 daemons up osd: 193 osds: 191 up (since 8m), 191 in (since 9m); 267 remapped pgs rgw: 2 daemons active (1 hosts, 1 zones)
data: volumes: 1/1 healthy pools: 10 pools, 793 pgs objects: 257.08M objects, 292 TiB usage: 614 TiB used, 386 TiB / 1000 TiB avail pgs: 835306/1285383707 objects degraded (0.065%) 225512620/1285383707 objects misplaced (17.544%) 525 active+clean 230 active+remapped+backfilling 32 active+remapped+backfill_toofull 5 active+undersized+degraded+remapped+backfilling 1 active+recovery_toofull+degraded
io: recovery: 978 MiB/s, 825 objects/s
---
Do not know if it is related but the cluster has been rebalancing for a few days now, after I've set EC pool only to use hdd.
---
If I start rgw with debug I get something like this in logs: [root@ctplmon2 ~]# radosgw -c /etc/ceph/ceph.conf --setuser ceph --setgroup ceph -n client.radosgw.moja.shramba.ctplmon2 -f -m 194.249.4.104:6789 --debug-rgw=99/99 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
In logs I get: 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 deferred set uid:gid to 167:167 (ceph:ceph) 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable), process radosgw, pid 168935 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework: beast 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework conf key: port, val: 4444 2024-12-21T23:16:59.898+0100 7f65a19257c0 1 radosgw_Main not setting numa affinity 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 rgw_d3n: rgw_d3n_l1_local_datacache_enabled=0 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 D3N datacache enabled: 0 2024-12-21T23:16:59.901+0100 7f658dffb640 20 reqs_thread_entry: start 2024-12-21T23:16:59.901+0100 7f658d7fa640 10 entry start 2024-12-21T23:16:59.908+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: realm 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 4 rgw main: RGWPeriod::init failed to init realm id : (2) No such file or directory 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: searching for the correct realm 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.215+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 zone default found 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Realm: () 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: ZoneGroup: default (b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578) 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Zone: default (c2c70444-7a41-4acd-a0d0-9f87d324ec72) 2024-12-21T23:17:00.295+0100 7f65a19257c0 10 cannot find current period zonegroup using local zonegroup configuration 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: zonegroup default 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started sync module instance, tier type = 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started zone id=c2c70444-7a41-4acd-a0d0-9f87d324ec72 (name=default) with tier type = 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
---
Any ideas what might cause rgw to stop working?
Kind regards, Rok _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
You are right again. Thank you. --- However is this really the right way (for all IO to stop) since cluster have enough capacity to rebalance? Why does not rebalance algorithm prevent one osd to be "too full"? Rok On Sun, Dec 22, 2024 at 12:00 AM Eugen Block <eblock@nde.ag> wrote:
The full OSD is most likely the reason. You can temporarily increase the threshold to 0.97 or so, but you need to prevent that to happen. The cluster usually starts warning you at 85%.
Zitat von Rok Jaklič <rjaklic@gmail.com>:
Hi,
for some reason radosgw stopped working.
Cluster status: [root@ctplmon1 ~]# ceph -v ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable) [root@ctplmon1 ~]# ceph -s cluster: id: 0a6e5422-ac75-4093-af20-528ee00cc847 health: HEALTH_ERR 6 OSD(s) experiencing slow operations in BlueStore 2 backfillfull osd(s) 1 full osd(s) 1 nearfull osd(s) Low space hindering backfill (add storage if this doesn't resolve itself): 32 pgs backfill_toofull Degraded data redundancy: 835306/1285383707 objects degraded (0.065%), 6 pgs degraded, 5 pgs undersized 76 pgs not deep-scrubbed in time 45 pgs not scrubbed in time Full OSDs blocking recovery: 1 pg recovery_toofull 9 pool(s) full 9 daemons have recently crashed
services: mon: 3 daemons, quorum ctplmon1,ctplmon3,ctplmon2 (age 36m) mgr: ctplmon1(active, since 65m) mds: 1/1 daemons up osd: 193 osds: 191 up (since 8m), 191 in (since 9m); 267 remapped pgs rgw: 2 daemons active (1 hosts, 1 zones)
data: volumes: 1/1 healthy pools: 10 pools, 793 pgs objects: 257.08M objects, 292 TiB usage: 614 TiB used, 386 TiB / 1000 TiB avail pgs: 835306/1285383707 objects degraded (0.065%) 225512620/1285383707 objects misplaced (17.544%) 525 active+clean 230 active+remapped+backfilling 32 active+remapped+backfill_toofull 5 active+undersized+degraded+remapped+backfilling 1 active+recovery_toofull+degraded
io: recovery: 978 MiB/s, 825 objects/s
---
Do not know if it is related but the cluster has been rebalancing for a few days now, after I've set EC pool only to use hdd.
---
If I start rgw with debug I get something like this in logs: [root@ctplmon2 ~]# radosgw -c /etc/ceph/ceph.conf --setuser ceph --setgroup ceph -n client.radosgw.moja.shramba.ctplmon2 -f -m 194.249.4.104:6789 --debug-rgw=99/99 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
In logs I get: 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 deferred set uid:gid to 167:167 (ceph:ceph) 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable), process radosgw, pid 168935 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework: beast 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework conf key: port, val: 4444 2024-12-21T23:16:59.898+0100 7f65a19257c0 1 radosgw_Main not setting numa affinity 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 rgw_d3n: rgw_d3n_l1_local_datacache_enabled=0 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 D3N datacache enabled: 0 2024-12-21T23:16:59.901+0100 7f658dffb640 20 reqs_thread_entry: start 2024-12-21T23:16:59.901+0100 7f658d7fa640 10 entry start 2024-12-21T23:16:59.908+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: realm 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 4 rgw main: RGWPeriod::init failed to init realm id : (2) No such file or directory 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: searching for the correct realm 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.215+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 zone default found 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Realm: () 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: ZoneGroup: default (b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578) 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Zone: default (c2c70444-7a41-4acd-a0d0-9f87d324ec72) 2024-12-21T23:17:00.295+0100 7f65a19257c0 10 cannot find current period zonegroup using local zonegroup configuration 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: zonegroup default 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started sync module instance, tier type = 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started zone id=c2c70444-7a41-4acd-a0d0-9f87d324ec72 (name=default) with tier type = 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
---
Any ideas what might cause rgw to stop working?
Kind regards, Rok _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Rok, The full_osd state is when IO is stopped, to protect the data from corruption. The 95% limit can be overshot by 1-2%, be careful when you increase the full_osd limit. The nearfull_osd limit warns at 90% and backfill onto the OSD is halted. But PGs still move off the OSD, usually these warnings resolve themselves during data move. If it doesn't after the upgrade then ofc you need to look. There are a couple of circumstances that might fill up an OSD. Poor balance, low number of PGs (top of my head). You could reweight the OSDs in question to tell Ceph to move some PGs off. But it leaves the PG move to the algorithm. You could also use the pgremapper to manually reassign PGs to different OSDs. This gives you more control over PG movement. This works by setting upmaps, the balancer needs to be off and the ceph version needs to be throughout newer than Luminous. https://github.com/digitalocean/pgremapper I hope this helps. Cheers, Alwin Antreich croit GmbH, https://croit.io/ On Sun, Dec 22, 2024, 00:07 Rok Jaklič <rjaklic@gmail.com> wrote:
You are right again.
Thank you.
---
However is this really the right way (for all IO to stop) since cluster have enough capacity to rebalance?
Why does not rebalance algorithm prevent one osd to be "too full"?
Rok
On Sun, Dec 22, 2024 at 12:00 AM Eugen Block <eblock@nde.ag> wrote:
The full OSD is most likely the reason. You can temporarily increase the threshold to 0.97 or so, but you need to prevent that to happen. The cluster usually starts warning you at 85%.
Zitat von Rok Jaklič <rjaklic@gmail.com>:
Hi,
for some reason radosgw stopped working.
Cluster status: [root@ctplmon1 ~]# ceph -v ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable) [root@ctplmon1 ~]# ceph -s cluster: id: 0a6e5422-ac75-4093-af20-528ee00cc847 health: HEALTH_ERR 6 OSD(s) experiencing slow operations in BlueStore 2 backfillfull osd(s) 1 full osd(s) 1 nearfull osd(s) Low space hindering backfill (add storage if this doesn't resolve itself): 32 pgs backfill_toofull Degraded data redundancy: 835306/1285383707 objects degraded (0.065%), 6 pgs degraded, 5 pgs undersized 76 pgs not deep-scrubbed in time 45 pgs not scrubbed in time Full OSDs blocking recovery: 1 pg recovery_toofull 9 pool(s) full 9 daemons have recently crashed
services: mon: 3 daemons, quorum ctplmon1,ctplmon3,ctplmon2 (age 36m) mgr: ctplmon1(active, since 65m) mds: 1/1 daemons up osd: 193 osds: 191 up (since 8m), 191 in (since 9m); 267 remapped pgs rgw: 2 daemons active (1 hosts, 1 zones)
data: volumes: 1/1 healthy pools: 10 pools, 793 pgs objects: 257.08M objects, 292 TiB usage: 614 TiB used, 386 TiB / 1000 TiB avail pgs: 835306/1285383707 objects degraded (0.065%) 225512620/1285383707 objects misplaced (17.544%) 525 active+clean 230 active+remapped+backfilling 32 active+remapped+backfill_toofull 5 active+undersized+degraded+remapped+backfilling 1 active+recovery_toofull+degraded
io: recovery: 978 MiB/s, 825 objects/s
---
Do not know if it is related but the cluster has been rebalancing for a few days now, after I've set EC pool only to use hdd.
---
If I start rgw with debug I get something like this in logs: [root@ctplmon2 ~]# radosgw -c /etc/ceph/ceph.conf --setuser ceph --setgroup ceph -n client.radosgw.moja.shramba.ctplmon2 -f -m 194.249.4.104:6789 --debug-rgw=99/99 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
In logs I get: 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 deferred set uid:gid to 167:167 (ceph:ceph) 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable), process radosgw, pid 168935 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework: beast 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework conf key: port, val: 4444 2024-12-21T23:16:59.898+0100 7f65a19257c0 1 radosgw_Main not setting numa affinity 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 rgw_d3n: rgw_d3n_l1_local_datacache_enabled=0 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 D3N datacache enabled: 0 2024-12-21T23:16:59.901+0100 7f658dffb640 20 reqs_thread_entry: start 2024-12-21T23:16:59.901+0100 7f658d7fa640 10 entry start 2024-12-21T23:16:59.908+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: realm 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 4 rgw main: RGWPeriod::init failed to init realm id : (2) No such file or directory 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: searching for the correct realm 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.215+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 zone default found 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Realm: () 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: ZoneGroup: default (b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578) 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Zone: default (c2c70444-7a41-4acd-a0d0-9f87d324ec72) 2024-12-21T23:17:00.295+0100 7f65a19257c0 10 cannot find current period zonegroup using local zonegroup configuration 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: zonegroup default 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started sync module instance, tier type = 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started zone id=c2c70444-7a41-4acd-a0d0-9f87d324ec72 (name=default) with tier type = 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
---
Any ideas what might cause rgw to stop working?
Kind regards, Rok _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Rok, All great suggestions here, try moving around some pgs with upmap. In case you need something very basic and simple we use this: https://github.com/laimis9133/plankton-swarm I'll add a target osd option this evening. Best, Laimis J. On Sun, Dec 22, 2024, 10:49 Alwin Antreich <alwin.antreich@croit.io> wrote:
Hi Rok,
The full_osd state is when IO is stopped, to protect the data from corruption. The 95% limit can be overshot by 1-2%, be careful when you increase the full_osd limit.
The nearfull_osd limit warns at 90% and backfill onto the OSD is halted. But PGs still move off the OSD, usually these warnings resolve themselves during data move. If it doesn't after the upgrade then ofc you need to look.
There are a couple of circumstances that might fill up an OSD. Poor balance, low number of PGs (top of my head).
You could reweight the OSDs in question to tell Ceph to move some PGs off. But it leaves the PG move to the algorithm.
You could also use the pgremapper to manually reassign PGs to different OSDs. This gives you more control over PG movement. This works by setting upmaps, the balancer needs to be off and the ceph version needs to be throughout newer than Luminous. https://github.com/digitalocean/pgremapper
I hope this helps.
Cheers, Alwin Antreich croit GmbH, https://croit.io/
On Sun, Dec 22, 2024, 00:07 Rok Jaklič <rjaklic@gmail.com> wrote:
You are right again.
Thank you.
---
However is this really the right way (for all IO to stop) since cluster have enough capacity to rebalance?
Why does not rebalance algorithm prevent one osd to be "too full"?
Rok
On Sun, Dec 22, 2024 at 12:00 AM Eugen Block <eblock@nde.ag> wrote:
The full OSD is most likely the reason. You can temporarily increase the threshold to 0.97 or so, but you need to prevent that to happen. The cluster usually starts warning you at 85%.
Zitat von Rok Jaklič <rjaklic@gmail.com>:
Hi,
for some reason radosgw stopped working.
Cluster status: [root@ctplmon1 ~]# ceph -v ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable) [root@ctplmon1 ~]# ceph -s cluster: id: 0a6e5422-ac75-4093-af20-528ee00cc847 health: HEALTH_ERR 6 OSD(s) experiencing slow operations in BlueStore 2 backfillfull osd(s) 1 full osd(s) 1 nearfull osd(s) Low space hindering backfill (add storage if this doesn't resolve itself): 32 pgs backfill_toofull Degraded data redundancy: 835306/1285383707 objects degraded (0.065%), 6 pgs degraded, 5 pgs undersized 76 pgs not deep-scrubbed in time 45 pgs not scrubbed in time Full OSDs blocking recovery: 1 pg recovery_toofull 9 pool(s) full 9 daemons have recently crashed
services: mon: 3 daemons, quorum ctplmon1,ctplmon3,ctplmon2 (age 36m) mgr: ctplmon1(active, since 65m) mds: 1/1 daemons up osd: 193 osds: 191 up (since 8m), 191 in (since 9m); 267 remapped pgs rgw: 2 daemons active (1 hosts, 1 zones)
data: volumes: 1/1 healthy pools: 10 pools, 793 pgs objects: 257.08M objects, 292 TiB usage: 614 TiB used, 386 TiB / 1000 TiB avail pgs: 835306/1285383707 objects degraded (0.065%) 225512620/1285383707 objects misplaced (17.544%) 525 active+clean 230 active+remapped+backfilling 32 active+remapped+backfill_toofull 5 active+undersized+degraded+remapped+backfilling 1 active+recovery_toofull+degraded
io: recovery: 978 MiB/s, 825 objects/s
---
Do not know if it is related but the cluster has been rebalancing for a few days now, after I've set EC pool only to use hdd.
---
If I start rgw with debug I get something like this in logs: [root@ctplmon2 ~]# radosgw -c /etc/ceph/ceph.conf --setuser ceph --setgroup ceph -n client.radosgw.moja.shramba.ctplmon2 -f -m 194.249.4.104:6789 --debug-rgw=99/99 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
In logs I get: 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 deferred set uid:gid to 167:167 (ceph:ceph) 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable), process radosgw, pid 168935 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework: beast 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework conf key: port, val: 4444 2024-12-21T23:16:59.898+0100 7f65a19257c0 1 radosgw_Main not setting numa affinity 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 rgw_d3n: rgw_d3n_l1_local_datacache_enabled=0 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 D3N datacache enabled: 0 2024-12-21T23:16:59.901+0100 7f658dffb640 20 reqs_thread_entry: start 2024-12-21T23:16:59.901+0100 7f658d7fa640 10 entry start 2024-12-21T23:16:59.908+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: realm 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 4 rgw main: RGWPeriod::init failed to init realm id : (2) No such file or directory 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: searching for the correct realm 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.215+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 zone default found 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Realm: () 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: ZoneGroup: default (b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578) 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Zone: default (c2c70444-7a41-4acd-a0d0-9f87d324ec72) 2024-12-21T23:17:00.295+0100 7f65a19257c0 10 cannot find current period zonegroup using local zonegroup configuration 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: zonegroup default 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started sync module instance, tier type = 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started zone id=c2c70444-7a41-4acd-a0d0-9f87d324ec72 (name=default) with tier type = 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
---
Any ideas what might cause rgw to stop working?
Kind regards, Rok _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Thank you all for your suggestions. I've increased full ratio to 0.96 and rgw started to work again. However: I've tried to set for e.g. osd.122 reweight, crush reweight and then also with: ceph osd pg-upmap-items 9.169 122 187 and then also [root@ctplmon1 plankton-swarm]# bash ./plankton-swarm.sh source-osds osd.122 3 Using custom source OSDs: osd.122 Underused OSDs (<65%): 63,108,200,199,143,144,198,142,140,146,195,141,148,194,196,197,103,158,191,147,125,25,164,54,193,126,192,145,46,19,68,157,128,15,60,53,134,129,87,0,102,131,6,43,78,127,120,33,81,119,3,5,74,190,70,8,160,24,58,156,114,29,186,82,96,116,182,48,84,28,18,44,178,39,4,75,115,76,79,130,72,86,159,40,184,22,26,35,71,171,88,64,175,187,170,165,41,110,94,150,111,17,83,14,27,49,2,37,124,172,177,98,152,104,95,118,168,189,132,52,105,32,59,21,101,9,107 Will now find ways to move 3 pgs in each OSD respecting node failure domain. Processing OSD osd.122... dumped all No active and clean pgs found for osd.122, skipping. Balance pgs commands written to swarm-file - review and let planktons swarm with 'bash swarm-file'. It seems it cannot find available osd. And utilization is still increasing. Is there any other method I can "force" that osd is being used for further balancing? Does the balancer need to be turned off for this command to work? Does degraded objects (even though a really small number 988069/1282375690 objects degraded (0.077%)) prevent ceph to properly balance data? Right now I am adding 2, 3 new osds every few hours, but it does not seem to help. --- Let us say 1 OSD gets full again. What if I "destroy that osd" (since we have enough other "free" osds I think)? Rok On Sun, Dec 22, 2024 at 10:42 AM Laimis Juzeliūnas < laimis.juzeliunas@oxylabs.io> wrote:
Hi Rok,
All great suggestions here, try moving around some pgs with upmap. In case you need something very basic and simple we use this: https://github.com/laimis9133/plankton-swarm
I'll add a target osd option this evening.
Best, Laimis J.
On Sun, Dec 22, 2024, 10:49 Alwin Antreich <alwin.antreich@croit.io> wrote:
Hi Rok,
The full_osd state is when IO is stopped, to protect the data from corruption. The 95% limit can be overshot by 1-2%, be careful when you increase the full_osd limit.
The nearfull_osd limit warns at 90% and backfill onto the OSD is halted. But PGs still move off the OSD, usually these warnings resolve themselves during data move. If it doesn't after the upgrade then ofc you need to look.
There are a couple of circumstances that might fill up an OSD. Poor balance, low number of PGs (top of my head).
You could reweight the OSDs in question to tell Ceph to move some PGs off. But it leaves the PG move to the algorithm.
You could also use the pgremapper to manually reassign PGs to different OSDs. This gives you more control over PG movement. This works by setting upmaps, the balancer needs to be off and the ceph version needs to be throughout newer than Luminous. https://github.com/digitalocean/pgremapper
I hope this helps.
Cheers, Alwin Antreich croit GmbH, https://croit.io/
On Sun, Dec 22, 2024, 00:07 Rok Jaklič <rjaklic@gmail.com> wrote:
You are right again.
Thank you.
---
However is this really the right way (for all IO to stop) since cluster have enough capacity to rebalance?
Why does not rebalance algorithm prevent one osd to be "too full"?
Rok
On Sun, Dec 22, 2024 at 12:00 AM Eugen Block <eblock@nde.ag> wrote:
The full OSD is most likely the reason. You can temporarily increase the threshold to 0.97 or so, but you need to prevent that to happen. The cluster usually starts warning you at 85%.
Zitat von Rok Jaklič <rjaklic@gmail.com>:
Hi,
for some reason radosgw stopped working.
Cluster status: [root@ctplmon1 ~]# ceph -v ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable) [root@ctplmon1 ~]# ceph -s cluster: id: 0a6e5422-ac75-4093-af20-528ee00cc847 health: HEALTH_ERR 6 OSD(s) experiencing slow operations in BlueStore 2 backfillfull osd(s) 1 full osd(s) 1 nearfull osd(s) Low space hindering backfill (add storage if this doesn't resolve itself): 32 pgs backfill_toofull Degraded data redundancy: 835306/1285383707 objects degraded (0.065%), 6 pgs degraded, 5 pgs undersized 76 pgs not deep-scrubbed in time 45 pgs not scrubbed in time Full OSDs blocking recovery: 1 pg recovery_toofull 9 pool(s) full 9 daemons have recently crashed
services: mon: 3 daemons, quorum ctplmon1,ctplmon3,ctplmon2 (age 36m) mgr: ctplmon1(active, since 65m) mds: 1/1 daemons up osd: 193 osds: 191 up (since 8m), 191 in (since 9m); 267 remapped pgs rgw: 2 daemons active (1 hosts, 1 zones)
data: volumes: 1/1 healthy pools: 10 pools, 793 pgs objects: 257.08M objects, 292 TiB usage: 614 TiB used, 386 TiB / 1000 TiB avail pgs: 835306/1285383707 objects degraded (0.065%) 225512620/1285383707 objects misplaced (17.544%) 525 active+clean 230 active+remapped+backfilling 32 active+remapped+backfill_toofull 5 active+undersized+degraded+remapped+backfilling 1 active+recovery_toofull+degraded
io: recovery: 978 MiB/s, 825 objects/s
---
Do not know if it is related but the cluster has been rebalancing for a few days now, after I've set EC pool only to use hdd.
---
If I start rgw with debug I get something like this in logs: [root@ctplmon2 ~]# radosgw -c /etc/ceph/ceph.conf --setuser ceph --setgroup ceph -n client.radosgw.moja.shramba.ctplmon2 -f -m 194.249.4.104:6789 --debug-rgw=99/99 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
In logs I get: 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 deferred set uid:gid to 167:167 (ceph:ceph) 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 ceph version 17.2.8 (f817ceb7f187defb1d021d6328fa833eb8e943b3) quincy (stable), process radosgw, pid 168935 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework: beast 2024-12-21T23:16:59.898+0100 7f65a19257c0 0 framework conf key: port, val: 4444 2024-12-21T23:16:59.898+0100 7f65a19257c0 1 radosgw_Main not setting numa affinity 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 rgw_d3n: rgw_d3n_l1_local_datacache_enabled=0 2024-12-21T23:16:59.901+0100 7f65a19257c0 1 D3N datacache enabled: 0 2024-12-21T23:16:59.901+0100 7f658dffb640 20 reqs_thread_entry: start 2024-12-21T23:16:59.901+0100 7f658d7fa640 10 entry start 2024-12-21T23:16:59.908+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: realm 2024-12-21T23:16:59.914+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 4 rgw main: RGWPeriod::init failed to init realm id : (2) No such file or directory 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:16:59.915+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:16:59.917+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:16:59.945+0100 7f65a19257c0 20 rgw main: searching for the correct realm 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.210+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.211+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.212+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.213+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.214+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.215+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_info.c2c70444-7a41-4acd-a0d0-9f87d324ec72 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroup_info.b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zone_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: RGWRados::pool_iterate: got zonegroups_names.default 2024-12-21T23:17:00.284+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.285+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=873 2024-12-21T23:17:00.286+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=46 2024-12-21T23:17:00.287+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=0 bl.length=358 2024-12-21T23:17:00.293+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 zone default found 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Realm: () 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: ZoneGroup: default (b1e0d55c-f7cb-4e73-b1cb-6cffa1fd6578) 2024-12-21T23:17:00.295+0100 7f65a19257c0 4 rgw main: Zone: default (c2c70444-7a41-4acd-a0d0-9f87d324ec72) 2024-12-21T23:17:00.295+0100 7f65a19257c0 10 cannot find current period zonegroup using local zonegroup configuration 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: zonegroup default 2024-12-21T23:17:00.295+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.296+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.299+0100 7f65a19257c0 20 rgw main: rados->read ofs=0 len=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: rados_obj.operate() r=-2 bl.length=0 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started sync module instance, tier type = 2024-12-21T23:17:00.303+0100 7f65a19257c0 20 rgw main: started zone id=c2c70444-7a41-4acd-a0d0-9f87d324ec72 (name=default) with tier type = 2024-12-21T23:21:59.898+0100 7f659e380640 -1 Initialization timeout, failed to initialize
---
Any ideas what might cause rgw to stop working?
Kind regards, Rok _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Rok, Try running (122 instead of osd.122): ./plankton-swarm.sh source-osds 122 3 bash swarm-file Will have to work on the naming conventions, apologies. The pgremapper tool also will be able to help in this case. Best, Laimis J.
On 22 Dec 2024, at 17:07, Rok Jaklič <rjaklic@gmail.com> wrote:
No active and clean pgs found for osd.122, skipping.
Got the same output for osd.122 and 122: [root@ctplmon1 plankton-swarm]# bash ./plankton-swarm.sh source-osds 122 3 Using custom source OSDs: 122 Underused OSDs (<65%): 63,108,143,144,142,140,146,200,141,148,199,195,198,194,196,103,197,158,147,19,125,191,164,25,126,54,145,46,68,193,192,157,134,131,6,15,128,3,60,33,53,129,87,0,102,43,78,127,160,81,119,178,120,79,5,190,72,132,74,156,114,82,70,8,58,24,18,29,49,130,76,96,116,71,48,26,84,50,170,175,165,28,186,184,44,39,110,182,4,98,115,88,75,86,159,187,22,35,83,40,171,64,95,152,111,150,94,41,17,104,52,14,30,27,173,2,37,9,105,101,124,62,32,45,149,172,176,1,168,177,118,106,100,59 Will now find ways to move 3 pgs in each OSD respecting node failure domain. Processing OSD 122... dumped all No active and clean pgs found for 122, skipping. Balance pgs commands written to swarm-file - review and let planktons swarm with 'bash swarm-file'. --- Got [root@ctplmon1 plankton-swarm]# ceph pg dump | grep ",122]" 9.3c 501445 0 0 271415 0 625758213726 0 0 10071 0 10071 active+remapped+backfilling 2024-12-22T18:34:45.510544+0100 319600'7529350 319600:22988194 [132,7,155,95,181] 132 [84,171,11,95,122] 84 316281'7471014 2024-12-18T07:08:00.062519+0100 315831'7421752 2024-12-15T09:24:50.719535+0100 0 3835 queued for deep scrub 501028 0 9.16 500984 0 0 431251 0 626842925755 0 0 9152 0 9152 active+remapped+backfilling 2024-12-22T13:15:21.725290+0100 319600'7542047 319600:26275250 [161,2,90,76,187] 161 [161,2,90,76,122] 161 316161'7465208 2024-12-17T11:35:38.075158+0100 316161'7465208 2024-12-17T11:35:38.075158+0100 0 26367 queued for scrub 503474 0 9.0 501542 0 0 410101 0 627802839022 0 0 10076 0 10076 active+remapped+backfilling 2024-12-22T14:28:09.575782+0100 319600'7510533 319600:25805615 [150,109,38,67,21] 150 [150,109,38,67,122] 150 316283'7453888 2024-12-18T07:33:10.588541+0100 316283'7453888 2024-12-18T07:33:10.588541+0100 0 21089 queued for scrub 501183 0 9.10 501310 0 0 81473 0 626139264394 0 0 10071 0 10071 active+remapped+backfilling 2024-12-22T13:12:55.859829+0100 319600'7533517 319600:29592030 [121,183,59,160,93] 121 [157,40,71,177,122] 157 316247'7472839 2024-12-18T01:34:50.713697+0100 314908'7370591 2024-12-11T21:07:36.550701+0100 0 606 queued for deep scrub 501175 0 9.14c 502664 0 0 63089 0 631327664452 0 0 10012 0 10012 active+remapped+backfilling 2024-12-21T23:10:58.757342+0100 319600'7529494 319600:29438730 [174,51,192,12,122] 174 [174,51,183,12,122] 174 316298'7471325 2024-12-18T10:38:34.981914+0100 316298'7471325 2024-12-18T10:38:34.981914+0100 0 12758 queued for scrub 502128 0 9.169 500251 0 0 454708 0 624120095931 0 0 10081 0 10081 active+remapped+backfilling 2024-12-22T15:18:50.038950+0100 319600'7569821 319600:25614989 [99,150,69,10,187] 99 [99,150,69,10,122] 99 316069'7480721 2024-12-16T21:34:22.271783+0100 316069'7480721 2024-12-16T21:34:22.271783+0100 0 59482 queued for scrub 501140 0 Its backfilling 122, that is why there is no output for grep -P 'active\+clean(?!\+)' afterwards. Rok On Sun, Dec 22, 2024 at 6:45 PM Laimis Juzeliūnas < laimis.juzeliunas@oxylabs.io> wrote:
Hi Rok,
Try running (122 instead of osd.122): ./plankton-swarm.sh source-osds 122 3 bash swarm-file
Will have to work on the naming conventions, apologies. The pgremapper tool also will be able to help in this case.
Best, Laimis J.
On 22 Dec 2024, at 17:07, Rok Jaklič <rjaklic@gmail.com> wrote:
No active and clean pgs found for osd.122, skipping.
I see, correct plankton will only consider clean+active pgs. It looks like the pgs on your output are remapped+backfilling, but not degraded or undersized. In that case you can try cancelling the upmaps (and backfills) with './upmap-remapped.py | sh' and then retry movement plankton. It usually takes a few hits of the script to clean up all the backfills. Upmap-remapped can be found here: https://github.com/cernceph/ceph-scripts/blob/master/tools/upmap/upmap-remap... Best, Laimis J. On Sun, Dec 22, 2024, 20:00 Rok Jaklič <rjaklic@gmail.com> wrote:
Got the same output for osd.122 and 122: [root@ctplmon1 plankton-swarm]# bash ./plankton-swarm.sh source-osds 122 3 Using custom source OSDs: 122 Underused OSDs (<65%): 63,108,143,144,142,140,146,200,141,148,199,195,198,194,196,103,197,158,147,19,125,191,164,25,126,54,145,46,68,193,192,157,134,131,6,15,128,3,60,33,53,129,87,0,102,43,78,127,160,81,119,178,120,79,5,190,72,132,74,156,114,82,70,8,58,24,18,29,49,130,76,96,116,71,48,26,84,50,170,175,165,28,186,184,44,39,110,182,4,98,115,88,75,86,159,187,22,35,83,40,171,64,95,152,111,150,94,41,17,104,52,14,30,27,173,2,37,9,105,101,124,62,32,45,149,172,176,1,168,177,118,106,100,59 Will now find ways to move 3 pgs in each OSD respecting node failure domain. Processing OSD 122... dumped all No active and clean pgs found for 122, skipping. Balance pgs commands written to swarm-file - review and let planktons swarm with 'bash swarm-file'.
---
Got
[root@ctplmon1 plankton-swarm]# ceph pg dump | grep ",122]" 9.3c 501445 0 0 271415 0 625758213726 0 0 10071 0 10071 active+remapped+backfilling 2024-12-22T18:34:45.510544+0100 319600'7529350 319600:22988194 [132,7,155,95,181] 132 [84,171,11,95,122] 84 316281'7471014 2024-12-18T07:08:00.062519+0100 315831'7421752 2024-12-15T09:24:50.719535+0100 0 3835 queued for deep scrub 501028 0 9.16 500984 0 0 431251 0 626842925755 0 0 9152 0 9152 active+remapped+backfilling 2024-12-22T13:15:21.725290+0100 319600'7542047 319600:26275250 [161,2,90,76,187] 161 [161,2,90,76,122] 161 316161'7465208 2024-12-17T11:35:38.075158+0100 316161'7465208 2024-12-17T11:35:38.075158+0100 0 26367 queued for scrub 503474 0 9.0 501542 0 0 410101 0 627802839022 0 0 10076 0 10076 active+remapped+backfilling 2024-12-22T14:28:09.575782+0100 319600'7510533 319600:25805615 [150,109,38,67,21] 150 [150,109,38,67,122] 150 316283'7453888 2024-12-18T07:33:10.588541+0100 316283'7453888 2024-12-18T07:33:10.588541+0100 0 21089 queued for scrub 501183 0 9.10 501310 0 0 81473 0 626139264394 0 0 10071 0 10071 active+remapped+backfilling 2024-12-22T13:12:55.859829+0100 319600'7533517 319600:29592030 [121,183,59,160,93] 121 [157,40,71,177,122] 157 316247'7472839 2024-12-18T01:34:50.713697+0100 314908'7370591 2024-12-11T21:07:36.550701+0100 0 606 queued for deep scrub 501175 0 9.14c 502664 0 0 63089 0 631327664452 0 0 10012 0 10012 active+remapped+backfilling 2024-12-21T23:10:58.757342+0100 319600'7529494 319600:29438730 [174,51,192,12,122] 174 [174,51,183,12,122] 174 316298'7471325 2024-12-18T10:38:34.981914+0100 316298'7471325 2024-12-18T10:38:34.981914+0100 0 12758 queued for scrub 502128 0 9.169 500251 0 0 454708 0 624120095931 0 0 10081 0 10081 active+remapped+backfilling 2024-12-22T15:18:50.038950+0100 319600'7569821 319600:25614989 [99,150,69,10,187] 99 [99,150,69,10,122] 99 316069'7480721 2024-12-16T21:34:22.271783+0100 316069'7480721 2024-12-16T21:34:22.271783+0100 0 59482 queued for scrub 501140 0
Its backfilling 122, that is why there is no output for grep -P 'active\+clean(?!\+)' afterwards.
Rok
On Sun, Dec 22, 2024 at 6:45 PM Laimis Juzeliūnas < laimis.juzeliunas@oxylabs.io> wrote:
Hi Rok,
Try running (122 instead of osd.122): ./plankton-swarm.sh source-osds 122 3 bash swarm-file
Will have to work on the naming conventions, apologies. The pgremapper tool also will be able to help in this case.
Best, Laimis J.
On 22 Dec 2024, at 17:07, Rok Jaklič <rjaklic@gmail.com> wrote:
No active and clean pgs found for osd.122, skipping.
Hi Rok, On Sun, 22 Dec 2024 at 16:08, Rok Jaklič <rjaklic@gmail.com> wrote:
Thank you all for your suggestions.
I've increased full ratio to 0.96 and rgw started to work again.
However:
I've tried to set for e.g. osd.122 reweight, crush reweight and then also with: ceph osd pg-upmap-items 9.169 122 187
A little too much at once. ;) 1) ceph osd reweight This overrules (lika a bias) the OSD weight temporarily but it will keep the crush weight of the bucket the same. Which means Ceph will try to keep the PGs within the node. 2) ceph osd crush reweight This sets the weight of the item in crush, which also tells Ceph that the node has a different weight and starts a new calculation for all the PGs connected to OSDs for this node. Which means more data movement. 3) ceph osd pg-upmap-items Overrides the placement of a PG specified by the crush algorithm. You can see these upmap entries in `ceph osd dump`. I do not recommend 1). With 3) you have the most control over where PGs should go. And it is best to keep the balancer off, otherwise it may interfere with your placements.
and then also [root@ctplmon1 plankton-swarm]# bash ./plankton-swarm.sh source-osds osd.122 3 Using custom source OSDs: osd.122 Underused OSDs (<65%): 63,108,200,199,143,144,198,142,140,146,195,141,148,194,196,197,103,158,191,147,125,25,164,54,193,126,192,145,46,19,68,157,128,15,60,53,134,129,87,0,102,131,6,43,78,127,120,33,81,119,3,5,74,190,70,8,160,24,58,156,114,29,186,82,96,116,182,48,84,28,18,44,178,39,4,75,115,76,79,130,72,86,159,40,184,22,26,35,71,171,88,64,175,187,170,165,41,110,94,150,111,17,83,14,27,49,2,37,124,172,177,98,152,104,95,118,168,189,132,52,105,32,59,21,101,9,107 Will now find ways to move 3 pgs in each OSD respecting node failure domain. Processing OSD osd.122... dumped all No active and clean pgs found for osd.122, skipping. Balance pgs commands written to swarm-file - review and let planktons swarm with 'bash swarm-file'.
On another post of yours you show 6x PGs being processed for osd.122, which is usually quite a lot for HDDs. Reducing the amount of osd_max_backfills (for wpq) to 1 will reduce the amount of parallel backfills. Where at most 2x backfills per OSD will happen. This might reduce the extra space needed during rebalance of these PGs, as the time needed to complete the rebalance of a particular PG will only increase the more PGs are processed in parallel. Once the PG has backfilled successfully it will be removed from the old OSD. You can also try to use `ceph pg force-backfill` of particular PGs to let Ceph know to process them prior to other PGs. Can you please post a `ceph osd df tree` and a `ceph df` to give us a picture on how things are? Cheers, Alwin croit GmbH, https://croit.io
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was most effective, but not for "all" osds I tried. Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on? Output of commands in attachment. Rok On Sun, Dec 22, 2024 at 7:41 PM Alwin Antreich <alwin.antreich@croit.io> wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 16:08, Rok Jaklič <rjaklic@gmail.com> wrote:
Thank you all for your suggestions.
I've increased full ratio to 0.96 and rgw started to work again.
However:
I've tried to set for e.g. osd.122 reweight, crush reweight and then also with: ceph osd pg-upmap-items 9.169 122 187
A little too much at once. ;)
1) ceph osd reweight This overrules (lika a bias) the OSD weight temporarily but it will keep the crush weight of the bucket the same. Which means Ceph will try to keep the PGs within the node.
2) ceph osd crush reweight This sets the weight of the item in crush, which also tells Ceph that the node has a different weight and starts a new calculation for all the PGs connected to OSDs for this node. Which means more data movement.
3) ceph osd pg-upmap-items Overrides the placement of a PG specified by the crush algorithm. You can see these upmap entries in `ceph osd dump`.
I do not recommend 1). With 3) you have the most control over where PGs should go. And it is best to keep the balancer off, otherwise it may interfere with your placements.
and then also [root@ctplmon1 plankton-swarm]# bash ./plankton-swarm.sh source-osds osd.122 3 Using custom source OSDs: osd.122 Underused OSDs (<65%): 63,108,200,199,143,144,198,142,140,146,195,141,148,194,196,197,103,158,191,147,125,25,164,54,193,126,192,145,46,19,68,157,128,15,60,53,134,129,87,0,102,131,6,43,78,127,120,33,81,119,3,5,74,190,70,8,160,24,58,156,114,29,186,82,96,116,182,48,84,28,18,44,178,39,4,75,115,76,79,130,72,86,159,40,184,22,26,35,71,171,88,64,175,187,170,165,41,110,94,150,111,17,83,14,27,49,2,37,124,172,177,98,152,104,95,118,168,189,132,52,105,32,59,21,101,9,107 Will now find ways to move 3 pgs in each OSD respecting node failure domain. Processing OSD osd.122... dumped all No active and clean pgs found for osd.122, skipping. Balance pgs commands written to swarm-file - review and let planktons swarm with 'bash swarm-file'.
On another post of yours you show 6x PGs being processed for osd.122, which is usually quite a lot for HDDs. Reducing the amount of osd_max_backfills (for wpq) to 1 will reduce the amount of parallel backfills. Where at most 2x backfills per OSD will happen. This might reduce the extra space needed during rebalance of these PGs, as the time needed to complete the rebalance of a particular PG will only increase the more PGs are processed in parallel. Once the PG has backfilled successfully it will be removed from the old OSD.
You can also try to use `ceph pg force-backfill` of particular PGs to let Ceph know to process them prior to other PGs.
Can you please post a `ceph osd df tree` and a `ceph df` to give us a picture on how things are?
Cheers, Alwin croit GmbH, https://croit.io
Hi Rok, On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was most effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts the usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account. I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement. Cheers, Alwin croit GmbH, https://croit.io/
Hi Alwin, Not to move too far away from the topic, but just wondering if there are any other recommendations regarding HDDs and backfills? We are currently doing in place node replacements and are quite on the aggressive side with our cluster when it comes to backfilling HDDs. So far we have seen target disks handle more than 20 parallel backfills with no issues. I’m just wondering if that’s pushing the limits too far. Rok, please let us know how things progress once the backfilling calms down. Very interested if you eventually get any luck with the plankton. Best, Laimis J.
On 22 Dec 2024, at 21:46, Alwin Antreich <alwin.antreich@croit.io> wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was most effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts the usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account.
I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement.
Cheers, Alwin croit GmbH, https://www.google.com/url?q=https://croit.io/&source=gmail-imap&ust=1735501671000000&usg=AOvVaw1rrBpyKfZiRg5DQjd0OCzn _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
autoscale_mode for pg is on for a particular pool (default.rgw.buckets.data) and EC 3-2 is used. During pool lifetime I've seen one time that PG number have changed automatically, but now I am also considering changing PG number manually after backfills completes. Right now pg_num 512 pgp_num 512 is used and I am considering to change it to 1024. Do you think that would be too aggressive maybe? Rok On Sun, Dec 22, 2024 at 8:46 PM Alwin Antreich <alwin.antreich@croit.io> wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was most effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts the usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account.
I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement.
Cheers, Alwin croit GmbH, https://croit.io/
However I now see that autoscaler is probably not working because of: ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.921+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool default.rgw.buckets.index won't scale due to overlapping roots: {-1, -18} ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.923+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool default.rgw.buckets.data won't scale due to overlapping roots: {-2, -1, -18} ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.929+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 1 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.929+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 2 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.930+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 3 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.931+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 4 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.931+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 5 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.932+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 6 contains an overlapping root -18... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.932+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 7 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.933+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 9 contains an overlapping root -2... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.934+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 10 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.934+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 11 contains an overlapping root -1... skipping scaling Rok On Mon, Dec 23, 2024 at 6:45 AM Rok Jaklič <rjaklic@gmail.com> wrote:
autoscale_mode for pg is on for a particular pool (default.rgw.buckets.data) and EC 3-2 is used. During pool lifetime I've seen one time that PG number have changed automatically, but now I am also considering changing PG number manually after backfills completes.
Right now pg_num 512 pgp_num 512 is used and I am considering to change it to 1024. Do you think that would be too aggressive maybe?
Rok
On Sun, Dec 22, 2024 at 8:46 PM Alwin Antreich <alwin.antreich@croit.io> wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was most effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts the usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account.
I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement.
Cheers, Alwin croit GmbH, https://croit.io/
Hi Rok, On Mon, 23 Dec 2024 at 07:28, Rok Jaklič <rjaklic@gmail.com> wrote:
However I now see that autoscaler is probably not working because of:
ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.921+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool default.rgw.buckets.index won't scale due to overlapping roots: {-1, -18} ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.923+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool default.rgw.buckets.data won't scale due to overlapping roots: {-2, -1, -18} ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.929+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 1 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.929+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 2 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.930+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 3 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.931+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 4 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.931+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 5 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.932+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 6 contains an overlapping root -18... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.932+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 7 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.933+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 9 contains an overlapping root -2... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.934+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 10 contains an overlapping root -1... skipping scaling ceph-mgr.ctplmon1.log:2024-12-23T07:12:00.934+0100 7f949edad640 0 [pg_autoscaler WARNING root] pool 11 contains an overlapping root -1... skipping scaling
This has been answered by Eugen in the other thread. ;)
Rok
On Mon, Dec 23, 2024 at 6:45 AM Rok Jaklič <rjaklic@gmail.com> wrote:
autoscale_mode for pg is on for a particular pool (default.rgw.buckets.data) and EC 3-2 is used. During pool lifetime I've seen one time that PG number have changed automatically, but now I am also considering changing PG number manually after backfills completes.
Right now pg_num 512 pgp_num 512 is used and I am considering to change it to 1024. Do you think that would be too aggressive maybe?
If my calculation is correct then your PGs are ~242 GB in size, which is
not bad. In general crush may distribute PGs better. It makes definitely sense to increase the number if you expect twice the amount of data to be stored. Cheers, Alwin croit GmbH, https://croit.io/
autoscale_mode for pg is on for a particular pool (default.rgw.buckets.data) and EC 3-2 is used. During pool lifetime I've seen one time that PG number have changed automatically
pg_num for a given pool likes to be a power of 2, so either the relative usage of pools or the overall cluster fillage has to change substantially for a change to be triggered in many cases.
but now I am also considering changing PG number manually after backfills completes.
If you do, be sure to disable the autoscaler for that pool.
Right now pg_num 512 pgp_num 512 is used and I am considering to change it to 1024. Do you think that would be too aggressive maybe?
Depends on how many OSDs you have and what the rest of the pools are like. Send us `ceph osd dump | grep pool` These days, assuming that your OSDs are BlueStore, chances are that going higher on pg_num won’t cause issues.
Rok
On Sun, Dec 22, 2024 at 8:46 PM Alwin Antreich <alwin.antreich@croit.io> wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was most effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts the usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account.
I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement.
Cheers, Alwin croit GmbH, https://croit.io/
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
[root@ctplmon1 ~]# ceph osd dump | grep pool pool 1 '.mgr' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 1 pgp_num 1 autoscale_mode on last_change 320144 flags hashpspool stripe_width 0 pg_num_min 1 application mgr,mgr_devicehealth pool 2 '.rgw.root' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/18964/18962 flags hashpspool stripe_width 0 application rgw pool 3 'default.rgw.log' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/127672/127670 flags hashpspool stripe_width 0 application rgw pool 4 'default.rgw.control' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/59850/59848 flags hashpspool stripe_width 0 application rgw pool 5 'default.rgw.meta' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 320144 lfor 0/51538/51536 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 8 application rgw pool 6 'default.rgw.buckets.index' replicated size 3 min_size 2 crush_rule 2 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 315285 lfor 0/127830/127828 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 8 application rgw pool 7 'default.rgw.buckets.non-ec' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/76474/76472 flags hashpspool stripe_width 0 application rgw pool 9 'default.rgw.buckets.data' erasure profile ec-32-profile size 5 min_size 4 crush_rule 1 object_hash rjenkins pg_num 512 pgp_num 512 autoscale_mode on last_change 320144 lfor 0/127784/214408 flags hashpspool,ec_overwrites stripe_width 12288 application rgw pool 10 'cephfs_data' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 128 pgp_num 128 autoscale_mode on last_change 320144 flags hashpspool,bulk stripe_width 0 application cephfs pool 11 'cephfs_metadata' replicated size 3 min_size 2 crush_rule 4 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 320144 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 16 recovery_priority 5 application cephfs --- Right now there are around 200 osds (5.5T) in a cluster, with around 25 waiting to be added. Rok On Mon, Dec 23, 2024 at 4:16 PM Anthony D'Atri <anthony.datri@gmail.com> wrote:
autoscale_mode for pg is on for a particular pool (default.rgw.buckets.data) and EC 3-2 is used. During pool lifetime I've seen one time that PG number have changed automatically
pg_num for a given pool likes to be a power of 2, so either the relative usage of pools or the overall cluster fillage has to change substantially for a change to be triggered in many cases.
but now I am also considering changing PG number manually after backfills completes.
If you do, be sure to disable the autoscaler for that pool.
Right now pg_num 512 pgp_num 512 is used and I am considering to change it to 1024. Do you think that would be too aggressive maybe?
Depends on how many OSDs you have and what the rest of the pools are like. Send us
`ceph osd dump | grep pool`
These days, assuming that your OSDs are BlueStore, chances are that going higher on pg_num won’t cause issues.
Rok
On Sun, Dec 22, 2024 at 8:46 PM Alwin Antreich <alwin.antreich@croit.io> wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was
effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts
most the
usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account.
I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement.
Cheers, Alwin croit GmbH, https://croit.io/
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
[root@ctplmon1 ~]# ceph osd dump | grep pool pool 1 '.mgr' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 1 pgp_num 1 autoscale_mode on last_change 320144 flags hashpspool stripe_width 0 pg_num_min 1 application mgr,mgr_devicehealth pool 2 '.rgw.root' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/18964/18962 flags hashpspool stripe_width 0 application rgw pool 3 'default.rgw.log' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/127672/127670 flags hashpspool stripe_width 0 application rgw pool 4 'default.rgw.control' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/59850/59848 flags hashpspool stripe_width 0 application rgw pool 5 'default.rgw.meta' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 320144 lfor 0/51538/51536 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 8 application rgw pool 6 'default.rgw.buckets.index' replicated size 3 min_size 2 crush_rule 2 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 315285 lfor 0/127830/127828 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 8 application rgw pool 7 'default.rgw.buckets.non-ec' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/76474/76472 flags hashpspool stripe_width 0 application rgw pool 9 'default.rgw.buckets.data' erasure profile ec-32-profile size 5 min_size 4 crush_rule 1 object_hash rjenkins pg_num 512 pgp_num 512 autoscale_mode on last_change 320144 lfor 0/127784/214408 flags hashpspool,ec_overwrites stripe_width 12288 application rgw pool 10 'cephfs_data' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 128 pgp_num 128 autoscale_mode on last_change 320144 flags hashpspool,bulk stripe_width 0 application cephfs pool 11 'cephfs_metadata' replicated size 3 min_size 2 crush_rule 4 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 320144 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 16 recovery_priority 5 application cephfs
---
Are you using HDDs, SSDs, or both? What does the PGs column at the right end of `ceph osd df` average? I’m still spinning up my brain this morning, but this seems reeeeeally low, like ~17 if all the OSDs are the same device class. buckets.index, notably, should be way higher. Assuming that your OSDs are all identical and thus that the index pool spans them all, I’d increase pg_num for the index pool and cephfs_metadata to 256 and for buckets.data to maybe 2048.
Right now there are around 200 osds (5.5T) in a cluster, with around 25 waiting to be added.
5.5T seems like an unusual number. Are these old HDDs, or perhaps 3DWPD SSDs?
Rok
On Mon, Dec 23, 2024 at 4:16 PM Anthony D'Atri <anthony.datri@gmail.com <mailto:anthony.datri@gmail.com>> wrote:
autoscale_mode for pg is on for a particular pool (default.rgw.buckets.data) and EC 3-2 is used. During pool lifetime I've seen one time that PG number have changed automatically
pg_num for a given pool likes to be a power of 2, so either the relative usage of pools or the overall cluster fillage has to change substantially for a change to be triggered in many cases.
but now I am also considering changing PG number manually after backfills completes.
If you do, be sure to disable the autoscaler for that pool.
Right now pg_num 512 pgp_num 512 is used and I am considering to change it to 1024. Do you think that would be too aggressive maybe?
Depends on how many OSDs you have and what the rest of the pools are like. Send us
`ceph osd dump | grep pool`
These days, assuming that your OSDs are BlueStore, chances are that going higher on pg_num won’t cause issues.
Rok
On Sun, Dec 22, 2024 at 8:46 PM Alwin Antreich <alwin.antreich@croit.io <mailto:alwin.antreich@croit.io>> wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com <mailto:rjaklic@gmail.com>> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was most effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts the usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account.
I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement.
Cheers, Alwin croit GmbH, https://croit.io/
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
For default.rgw.buckets.index ssd-s (actually nvme-s), and for default.rgw.buckets.data hdd-s. Average PG is around ~17.7. Actually most of the disks are Seagate 6T <https://www.amazon.com/Seagate-Enterprise-Capacity-ST6000NM0095-7200RPM/dp/B01CG0DBXE> in size, but this "translates" to 5.5T in ceph and yes, they are pretty old (from 2016, 2017 up to 2021). Rok On Mon, Dec 23, 2024 at 4:41 PM Anthony D'Atri <anthony.datri@gmail.com> wrote:
[root@ctplmon1 ~]# ceph osd dump | grep pool pool 1 '.mgr' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 1 pgp_num 1 autoscale_mode on last_change 320144 flags hashpspool stripe_width 0 pg_num_min 1 application mgr,mgr_devicehealth pool 2 '.rgw.root' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/18964/18962 flags hashpspool stripe_width 0 application rgw pool 3 'default.rgw.log' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/127672/127670 flags hashpspool stripe_width 0 application rgw pool 4 'default.rgw.control' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/59850/59848 flags hashpspool stripe_width 0 application rgw pool 5 'default.rgw.meta' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 320144 lfor 0/51538/51536 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 8 application rgw pool 6 'default.rgw.buckets.index' replicated size 3 min_size 2 crush_rule 2 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 315285 lfor 0/127830/127828 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 8 application rgw pool 7 'default.rgw.buckets.non-ec' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/76474/76472 flags hashpspool stripe_width 0 application rgw pool 9 'default.rgw.buckets.data' erasure profile ec-32-profile size 5 min_size 4 crush_rule 1 object_hash rjenkins pg_num 512 pgp_num 512 autoscale_mode on last_change 320144 lfor 0/127784/214408 flags hashpspool,ec_overwrites stripe_width 12288 application rgw pool 10 'cephfs_data' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 128 pgp_num 128 autoscale_mode on last_change 320144 flags hashpspool,bulk stripe_width 0 application cephfs pool 11 'cephfs_metadata' replicated size 3 min_size 2 crush_rule 4 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 320144 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 16 recovery_priority 5 application cephfs
---
Are you using HDDs, SSDs, or both? What does the PGs column at the right end of `ceph osd df` average? I’m still spinning up my brain this morning, but this seems reeeeeally low, like ~17 if all the OSDs are the same device class.
buckets.index, notably, should be way higher. Assuming that your OSDs are all identical and thus that the index pool spans them all, I’d increase pg_num for the index pool and cephfs_metadata to 256 and for buckets.data to maybe 2048.
Right now there are around 200 osds (5.5T) in a cluster, with around 25 waiting to be added.
5.5T seems like an unusual number. Are these old HDDs, or perhaps 3DWPD SSDs?
Rok
On Mon, Dec 23, 2024 at 4:16 PM Anthony D'Atri <anthony.datri@gmail.com> wrote:
autoscale_mode for pg is on for a particular pool (default.rgw.buckets.data) and EC 3-2 is used. During pool lifetime I've seen one time that PG number have changed automatically
pg_num for a given pool likes to be a power of 2, so either the relative usage of pools or the overall cluster fillage has to change substantially for a change to be triggered in many cases.
but now I am also considering changing PG number manually after backfills completes.
If you do, be sure to disable the autoscaler for that pool.
Right now pg_num 512 pgp_num 512 is used and I am considering to change it to 1024. Do you think that would be too aggressive maybe?
Depends on how many OSDs you have and what the rest of the pools are like. Send us
`ceph osd dump | grep pool`
These days, assuming that your OSDs are BlueStore, chances are that going higher on pg_num won’t cause issues.
Rok
On Sun, Dec 22, 2024 at 8:46 PM Alwin Antreich <alwin.antreich@croit.io
wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was
most
effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts the usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account.
I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement.
Cheers, Alwin croit GmbH, https://croit.io/
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On Dec 23, 2024, at 12:08 PM, Rok Jaklič <rjaklic@gmail.com> wrote:
For default.rgw.buckets.index ssd-s (actually nvme-s), and for default.rgw.buckets.data hdd-s. Average PG is around ~17.7.
Yikes, that is not doing you any favors at all, in terms of performance and of uniform OSD utilization. The party line is currently a target ratio of 100, though I have a PR open to return it to the former target of 200. I’d really like to make that 500 but we need to be somewhat conservative. Once you get your CRUSH rules / device classes sorted out the autoscaler should grow your pg_nums substantially, or you can take a walk on the wild side by turning it off and calculating yourself, old-school:-style https://docs.ceph.com/en/squid/rados/operations/pgcalc/
Actually most of the disks are Seagate 6T <https://www.amazon.com/Seagate-Enterprise-Capacity-ST6000NM0095-7200RPM/dp/B01CG0DBXE> in size, but this "translates" to 5.5T in ceph and yes, they are pretty old (from 2016, 2017 up to 2021).
Ack, I suspected so. That “translation” is in large part due to storage manufacturers being weasels: they describe devices in terms of base-10 units (TB), and humans and everyone else mainly think base-2 units (TiB). 6.0 TB = 5.45697 TiB. Back like 10-12 years ago Apple switched the macOS Finder from using the former to the latter and people were outraged because they believed that Apple was taking storage away.
Rok
On Mon, Dec 23, 2024 at 4:41 PM Anthony D'Atri <anthony.datri@gmail.com> wrote:
[root@ctplmon1 ~]# ceph osd dump | grep pool pool 1 '.mgr' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 1 pgp_num 1 autoscale_mode on last_change 320144 flags hashpspool stripe_width 0 pg_num_min 1 application mgr,mgr_devicehealth pool 2 '.rgw.root' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/18964/18962 flags hashpspool stripe_width 0 application rgw pool 3 'default.rgw.log' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/127672/127670 flags hashpspool stripe_width 0 application rgw pool 4 'default.rgw.control' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/59850/59848 flags hashpspool stripe_width 0 application rgw pool 5 'default.rgw.meta' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 320144 lfor 0/51538/51536 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 8 application rgw pool 6 'default.rgw.buckets.index' replicated size 3 min_size 2 crush_rule 2 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 315285 lfor 0/127830/127828 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 8 application rgw pool 7 'default.rgw.buckets.non-ec' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 32 pgp_num 32 autoscale_mode on last_change 320144 lfor 0/76474/76472 flags hashpspool stripe_width 0 application rgw pool 9 'default.rgw.buckets.data' erasure profile ec-32-profile size 5 min_size 4 crush_rule 1 object_hash rjenkins pg_num 512 pgp_num 512 autoscale_mode on last_change 320144 lfor 0/127784/214408 flags hashpspool,ec_overwrites stripe_width 12288 application rgw pool 10 'cephfs_data' replicated size 3 min_size 2 crush_rule 0 object_hash rjenkins pg_num 128 pgp_num 128 autoscale_mode on last_change 320144 flags hashpspool,bulk stripe_width 0 application cephfs pool 11 'cephfs_metadata' replicated size 3 min_size 2 crush_rule 4 object_hash rjenkins pg_num 8 pgp_num 8 autoscale_mode on last_change 320144 flags hashpspool stripe_width 0 pg_autoscale_bias 4 pg_num_min 16 recovery_priority 5 application cephfs
---
Are you using HDDs, SSDs, or both? What does the PGs column at the right end of `ceph osd df` average? I’m still spinning up my brain this morning, but this seems reeeeeally low, like ~17 if all the OSDs are the same device class.
buckets.index, notably, should be way higher. Assuming that your OSDs are all identical and thus that the index pool spans them all, I’d increase pg_num for the index pool and cephfs_metadata to 256 and for buckets.data to maybe 2048.
Right now there are around 200 osds (5.5T) in a cluster, with around 25 waiting to be added.
5.5T seems like an unusual number. Are these old HDDs, or perhaps 3DWPD SSDs?
Rok
On Mon, Dec 23, 2024 at 4:16 PM Anthony D'Atri <anthony.datri@gmail.com> wrote:
autoscale_mode for pg is on for a particular pool (default.rgw.buckets.data) and EC 3-2 is used. During pool lifetime I've seen one time that PG number have changed automatically
pg_num for a given pool likes to be a power of 2, so either the relative usage of pools or the overall cluster fillage has to change substantially for a change to be triggered in many cases.
but now I am also considering changing PG number manually after backfills completes.
If you do, be sure to disable the autoscaler for that pool.
Right now pg_num 512 pgp_num 512 is used and I am considering to change it to 1024. Do you think that would be too aggressive maybe?
Depends on how many OSDs you have and what the rest of the pools are like. Send us
`ceph osd dump | grep pool`
These days, assuming that your OSDs are BlueStore, chances are that going higher on pg_num won’t cause issues.
Rok
On Sun, Dec 22, 2024 at 8:46 PM Alwin Antreich <alwin.antreich@croit.io
wrote:
Hi Rok,
On Sun, 22 Dec 2024 at 20:19, Rok Jaklič <rjaklic@gmail.com> wrote:
First I tried with osd reweight, waited a few hours then osd crush reweight, then with pg-umpap from Laimis. Seems to crush reweight was
most
effective, but not for "all" osds I tried.
Uh, probably I've set ceph config set osd osd_max_backfills to high number in the past, probably better to reduce it to 1 in steps, since now much backfilling is already going on?
Every time a backfill finishes, a new one will be placed in the queue. The number of backfills won't reduce as long as you don't lower it. You can adjust it and see if it improves the backfill process or not (wait an hour or two).
Output of commands in attachment.
There seems to be a low amount of PGs for the rgw data pool, compared to the amount of OSDs. Though it depends on the EC profile and size of a shard (`ceph pg <id> query`) if this is really an issue. But in general the amount of PGs is important, because too few of them will make them grow larger. Hence backfilling a PG will take a longer time and easier tilts the usage of OSDs, as the algorithm works by pseudo-randomly placing PGs and not taking its size into account.
I'd wait with the PG adjustment after the backfilling to the HDDs has finished, should you need to adjust the number of PGs. As this will create more data movement.
Cheers, Alwin croit GmbH, https://croit.io/
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On 24/12/24 04:37, Anthony D'Atri wrote:
humans and everyone else mainly think base-2 units (TiB).
For my part, I see humans thinking in base-10 units. I'm glad we've at least got terms for the base-2 ones, so it can be clear in technical settings. At best however I see people having to clarify and getting frustrated with there being two options, one of which being less intuitive, particularly given that only in IT did we ever pretend a kilo-something was not a thousand. Therefore I'm on a collaborative campaign to standardise all of my org's thinking, talking and reporting to base-10. I'm not sure it will ever complete.
participants (6)
-
Alwin Antreich
-
Anthony D'Atri
-
Eugen Block
-
Gregory Orange
-
Laimis Juzeliūnas
-
Rok Jaklič