objects misplaced jumps up at 5%
Dear All, After adding 10 new nodes, each with 10 OSDs to a cluster, we are unable to get "objects misplaced" back to zero. The cluster successfully re-balanced from ~35% to 5% misplaced, however every time "objects misplaced" drops below 5%, a number of pgs start to backfill, increasing the "objects misplaced" to 5.1% I do not believe the balancer is active: [root@ceph7 ceph]# ceph balancer status { "last_optimize_duration": "", "plans": [], "mode": "upmap", "active": false, "optimize_result": "", "last_optimize_started": "" } The cluster has now been stuck at ~5% misplaced for a couple of weeks. The recovery is using ~1GiB/s bandwidth, and is preventing any scrubs. The cluster contains 2.6PB of cephfs, that is still read/write usable. Cluster originally had 10 nodes, each with 45 8TB drives. The new nodes have 10 x 16TB drives. To show the cluster before and immediately after an "episode" *************************************************** [root@ceph7 ceph]# ceph -s cluster: id: 36ed7113-080c-49b8-80e2-4947cc456f2a health: HEALTH_WARN 7 nearfull osd(s) 2 pool(s) nearfull Low space hindering backfill (add storage if this doesn't resolve itself): 11 pgs backfill_toofull 16372 pgs not deep-scrubbed in time 16372 pgs not scrubbed in time 1/3 mons down, quorum ceph1b,ceph3b services: mon: 3 daemons, quorum ceph1b,ceph3b (age 6d), out of quorum: ceph2b mgr: ceph3(active, since 3d), standbys: ceph1 mds: cephfs:1 {0=ceph1=up:active} 1 up:standby-replay osd: 554 osds: 554 up (since 4d), 554 in (since 5w); 848 remapped pgs task status: scrub status: mds.ceph1: idle mds.ceph2: idle data: pools: 3 pools, 16417 pgs objects: 937.39M objects, 2.6 PiB usage: 3.2 PiB used, 1.4 PiB / 4.6 PiB avail pgs: 467620187/9352502650 objects misplaced (5.000%) 7893 active+clean 7294 active+clean+snaptrim_wait 785 active+remapped+backfill_wait 382 active+clean+snaptrim 52 active+remapped+backfilling 11 active+remapped+backfill_wait+backfill_toofull io: client: 129 KiB/s rd, 82 MiB/s wr, 3 op/s rd, 53 op/s wr recovery: 1.1 GiB/s, 364 objects/s *************************************************** and then seconds later: *************************************************** [root@ceph7 ceph]# ceph -s cluster: id: 36ed7113-080c-49b8-80e2-4947cc456f2a health: HEALTH_WARN 7 nearfull osd(s) 2 pool(s) nearfull Low space hindering backfill (add storage if this doesn't resolve itself): 11 pgs backfill_toofull 16372 pgs not deep-scrubbed in time 16372 pgs not scrubbed in time 1/3 mons down, quorum ceph1b,ceph3b services: mon: 3 daemons, quorum ceph1b,ceph3b (age 6d), out of quorum: ceph2b mgr: ceph3(active, since 3d), standbys: ceph1 mds: cephfs:1 {0=ceph1=up:active} 1 up:standby-replay osd: 554 osds: 554 up (since 5d), 554 in (since 5w); 854 remapped pgs task status: scrub status: mds.ceph1: idle mds.ceph2: idle data: pools: 3 pools, 16417 pgs objects: 937.40M objects, 2.6 PiB usage: 3.2 PiB used, 1.4 PiB / 4.6 PiB avail pgs: 470821753/9352518510 objects misplaced (5.034%) 7892 active+clean 7290 active+clean+snaptrim_wait 791 active+remapped+backfill_wait 381 active+clean+snaptrim 52 active+remapped+backfilling 11 active+remapped+backfill_wait+backfill_toofull io: client: 155 KiB/s rd, 125 MiB/s wr, 2 op/s rd, 53 op/s wr recovery: 969 MiB/s, 330 objects/s *************************************************** If it helps, I've tried capturing 1/5 debug logs from an OSD. Not sure, but I think this is the way to follow a thread handling one pg as it decides to rebalance: [root@ceph7 ceph]# grep 7f2e569e9700 ceph-osd.312.log | less 2020-09-24 14:44:36.844 7f2e569e9700 1 osd.312 pg_epoch: 106808 pg[5.157ds0( v 106803'6043528 (103919'6040524,106803'6043528] local-lis/les=102671/102672 n=56293 ec=85890/1818 lis/c 102671/10 2671 les/c/f 102672/102672/0 106808/106808/106808) [148,508,398,457,256,533,137,469,357,306]p148(0) r=-1 lpr=106808 pi=[102671,106808)/1 luod=0'0 crt=106803'6043528 lcod 106801'6043526 active mbc={} ps=104] start_peering_interval up [312,424,369,461,546,525,498,169,251,127] -> [148,508,398,457,256,533,137,469,357,306], acting [312,424,369,461,546,525,498,169,251,127] -> [148,508,39 8,457,256,533,137,469,357,306], acting_primary 312(0) -> 148, up_primary 312(0) -> 148, role 0 -> -1, features acting 4611087854031667199 upacting 4611087854031667199 2020-09-24 14:44:36.847 7f2e569e9700 1 osd.312 pg_epoch: 106808 pg[5.157ds0( v 106803'6043528 (103919'6040524,106803'6043528] local-lis/les=102671/102672 n=56293 ec=85890/1818 lis/c 102671/102671 les/c/f 102672/102672/0 106808/106808/106808) [148,508,398,457,256,533,137,469,357,306]p148(0) r=-1 lpr=106808 pi=[102671,106808)/1 crt=106803'6043528 lcod 106801'6043526 unknown NOTIFY mbc={} ps=104] state<Start>: transitioning to Stray 2020-09-24 14:44:37.792 7f2e569e9700 1 osd.312 pg_epoch: 106809 pg[5.157ds0( v 106803'6043528 (103919'6040524,106803'6043528] local-lis/les=102671/102672 n=56293 ec=85890/1818 lis/c 102671/102671 les/c/f 102672/102672/0 106808/106809/106809) [148,508,398,457,256,533,137,469,357,306]/[312,424,369,461,546,525,498,169,251,127]p312(0) r=0 lpr=106809 pi=[102671,106809)/1 crt=106803'6043528 lcod 106801'6043526 mlcod 0'0 remapped NOTIFY mbc={} ps=104] start_peering_interval up [148,508,398,457,256,533,137,469,357,306] -> [148,508,398,457,256,533,137,469,357,306], acting [148,508,398,457,256,533,137,469,357,306] -> [312,424,369,461,546,525,498,169,251,127], acting_primary 148(0) -> 312, up_primary 148(0) -> 148, role -1 -> 0, features acting 4611087854031667199 upacting 4611087854031667199 2020-09-24 14:44:37.793 7f2e569e9700 1 osd.312 pg_epoch: 106809 pg[5.157ds0( v 106803'6043528 (103919'6040524,106803'6043528] local-lis/les=102671/102672 n=56293 ec=85890/1818 lis/c 102671/102671 les/c/f 102672/102672/0 106808/106809/106809) [148,508,398,457,256,533,137,469,357,306]/[312,424,369,461,546,525,498,169,251,127]p312(0) r=0 lpr=106809 pi=[102671,106809)/1 crt=106803'6043528 lcod 106801'6043526 mlcod 0'0 remapped mbc={} ps=104] state<Start>: transitioning to Primary 2020-09-24 14:44:38.832 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.137(6) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.861 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.148(0) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.879 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.256(4) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.894 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.306(9) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.902 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.357(8) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.912 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.398(2) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.923 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.457(3) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.931 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.469(7) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.938 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.508(1) from (0'0,0'0] MAX to 106803'6043528 2020-09-24 14:44:38.947 7f2e569e9700 0 log_channel(cluster) log [DBG] : 5.157ds0 starting backfill to osd.533(5) from (0'0,0'0] MAX to 106803'6043528 *************************************************** any advice appreciated, Jake -- Dr Jake Grimmett Head Of Scientific Computing MRC Laboratory of Molecular Biology Francis Crick Avenue, Cambridge CB2 0QH, UK.
On 2020-09-28 11:45, Jake Grimmett wrote:
To show the cluster before and immediately after an "episode"
***************************************************
[root@ceph7 ceph]# ceph -s cluster: id: 36ed7113-080c-49b8-80e2-4947cc456f2a health: HEALTH_WARN 7 nearfull osd(s) 2 pool(s) nearfull Low space hindering backfill (add storage if this doesn't resolve itself): 11 pgs backfill_toofull
What version are you running? I'm worried the nearfull OSDs might be the culprit here. There has been a bug with respect to neafull OSDs [1] that has been fixed since. You might or might not hit that. Check with "ceph osd df" to see if there are OSDs really too full or not. You can use Dan's upmap-remapped.py [2] to remap the PGs back to their original location and get the cluster in HEALTH_OK again. You might want to select deep-scrub by hand to make sure you get the most efficient way of deep-scrubbing (instead of randomly choosing a PG to deep-scrub). Gr. Stefan [1]: https://tracker.ceph.com/issues/39555 [2]: https://github.com/cernceph/ceph-scripts/blob/master/tools/upmap/upmap-remap...
Hi Stefan, many thanks for your good advice. We are using ceph version 14.2.11 There is an issue with full osds - I'm not sure it's causing this misplaced jump problem; I've reweighting the most full osds on several consecutive days to reduce the number of nearfull osds, and it seems to have no effect on the misplaced jump. I've not done a reweight for a few days, so we have a lot of very full osds. The osd balance is way off, so something is amis. Out of 550 OSD we see this spread in use (sorted least full to most full): ID WEIGHT REWEIGHT %USE VAR PGS 430 7.27730 1.00000 46.81 0.68 171 73 7.27730 1.00000 46.91 0.68 170 189 7.27730 1.00000 47.15 0.68 173 199 7.27730 1.00000 47.24 0.69 172 86 7.27730 1.00000 48.62 0.71 176 234 7.27730 1.00000 48.73 0.71 178 437 7.27730 1.00000 49.65 0.72 182 288 7.27730 1.00000 50.12 0.73 184 (SNIP) ID WEIGHT REWEIGHT %USE VAR PGS 455 14.55299 1.00000 84.39 1.23 619 541 14.55299 1.00000 84.40 1.23 620 456 14.55299 1.00000 84.73 1.23 621 487 14.55299 0.90002 85.56 1.24 620 527 14.55299 1.00000 86.61 1.26 638 466 14.55299 0.90002 86.78 1.26 639 501 14.55299 1.00000 87.39 1.27 645 542 14.55299 1.00000 88.06 1.28 645 462 14.55299 0.95001 91.23 1.32 670 549 14.55299 1.00000 91.45 1.33 676 I like your idea of remapping the pgs to their original location, and then re-balancing the osds to a sensible arrangement. will see if this works and report back.... best regards, Jake On 28/09/2020 11:08, Stefan Kooman wrote:
On 2020-09-28 11:45, Jake Grimmett wrote:
To show the cluster before and immediately after an "episode"
***************************************************
[root@ceph7 ceph]# ceph -s cluster: id: 36ed7113-080c-49b8-80e2-4947cc456f2a health: HEALTH_WARN 7 nearfull osd(s) 2 pool(s) nearfull Low space hindering backfill (add storage if this doesn't resolve itself): 11 pgs backfill_toofull
What version are you running? I'm worried the nearfull OSDs might be the culprit here. There has been a bug with respect to neafull OSDs [1] that has been fixed since. You might or might not hit that. Check with "ceph osd df" to see if there are OSDs really too full or not.
You can use Dan's upmap-remapped.py [2] to remap the PGs back to their original location and get the cluster in HEALTH_OK again. You might want to select deep-scrub by hand to make sure you get the most efficient way of deep-scrubbing (instead of randomly choosing a PG to deep-scrub).
Gr. Stefan
[1]: https://tracker.ceph.com/issues/39555 [2]: https://github.com/cernceph/ceph-scripts/blob/master/tools/upmap/upmap-remap...
-- Dr Jake Grimmett Head Of Scientific Computing MRC Laboratory of Molecular Biology Francis Crick Avenue, Cambridge CB2 0QH, UK.
Hi, 5% misplaced is the default target ratio for misplaced PGs when any automated rebalancing happens, the sources for this are either the balancer or pg scaling. So I'd suspect that there's a PG change ongoing (either pg autoscaler or a manual change, both obey the target misplaced ratio). You can check this by running "ceph osd pool ls detail" and check for the value of pg target. Also: Looks like you've set osd_scrub_during_recovery = false, this setting can be annoying on large erasure-coded setups on HDDs that see long recovery times. It's better to get IO priorities right; search mailing list for osd op queue cut off high. Paul On Mon, Sep 28, 2020 at 11:45 AM Jake Grimmett <jog@mrc-lmb.cam.ac.uk> wrote:
Dear All,
After adding 10 new nodes, each with 10 OSDs to a cluster, we are unable
to get "objects misplaced" back to zero.
The cluster successfully re-balanced from ~35% to 5% misplaced, however
every time "objects misplaced" drops below 5%, a number of pgs start to
backfill, increasing the "objects misplaced" to 5.1%
I do not believe the balancer is active:
[root@ceph7 ceph]# ceph balancer status
{
"last_optimize_duration": "",
"plans": [],
"mode": "upmap",
"active": false,
"optimize_result": "",
"last_optimize_started": ""
}
The cluster has now been stuck at ~5% misplaced for a couple of weeks.
The recovery is using ~1GiB/s bandwidth, and is preventing any scrubs.
The cluster contains 2.6PB of cephfs, that is still read/write usable.
Cluster originally had 10 nodes, each with 45 8TB drives. The new nodes
have 10 x 16TB drives.
To show the cluster before and immediately after an "episode"
***************************************************
[root@ceph7 ceph]# ceph -s
cluster:
id: 36ed7113-080c-49b8-80e2-4947cc456f2a
health: HEALTH_WARN
7 nearfull osd(s)
2 pool(s) nearfull
Low space hindering backfill (add storage if this doesn't
resolve itself): 11 pgs backfill_toofull
16372 pgs not deep-scrubbed in time
16372 pgs not scrubbed in time
1/3 mons down, quorum ceph1b,ceph3b
services:
mon: 3 daemons, quorum ceph1b,ceph3b (age 6d), out of quorum: ceph2b
mgr: ceph3(active, since 3d), standbys: ceph1
mds: cephfs:1 {0=ceph1=up:active} 1 up:standby-replay
osd: 554 osds: 554 up (since 4d), 554 in (since 5w); 848 remapped pgs
task status:
scrub status:
mds.ceph1: idle
mds.ceph2: idle
data:
pools: 3 pools, 16417 pgs
objects: 937.39M objects, 2.6 PiB
usage: 3.2 PiB used, 1.4 PiB / 4.6 PiB avail
pgs: 467620187/9352502650 objects misplaced (5.000%)
7893 active+clean
7294 active+clean+snaptrim_wait
785 active+remapped+backfill_wait
382 active+clean+snaptrim
52 active+remapped+backfilling
11 active+remapped+backfill_wait+backfill_toofull
io:
client: 129 KiB/s rd, 82 MiB/s wr, 3 op/s rd, 53 op/s wr
recovery: 1.1 GiB/s, 364 objects/s
***************************************************
and then seconds later:
***************************************************
[root@ceph7 ceph]# ceph -s
cluster:
id: 36ed7113-080c-49b8-80e2-4947cc456f2a
health: HEALTH_WARN
7 nearfull osd(s)
2 pool(s) nearfull
Low space hindering backfill (add storage if this doesn't
resolve itself): 11 pgs backfill_toofull
16372 pgs not deep-scrubbed in time
16372 pgs not scrubbed in time
1/3 mons down, quorum ceph1b,ceph3b
services:
mon: 3 daemons, quorum ceph1b,ceph3b (age 6d), out of quorum: ceph2b
mgr: ceph3(active, since 3d), standbys: ceph1
mds: cephfs:1 {0=ceph1=up:active} 1 up:standby-replay
osd: 554 osds: 554 up (since 5d), 554 in (since 5w); 854 remapped pgs
task status:
scrub status:
mds.ceph1: idle
mds.ceph2: idle
data:
pools: 3 pools, 16417 pgs
objects: 937.40M objects, 2.6 PiB
usage: 3.2 PiB used, 1.4 PiB / 4.6 PiB avail
pgs: 470821753/9352518510 objects misplaced (5.034%)
7892 active+clean
7290 active+clean+snaptrim_wait
791 active+remapped+backfill_wait
381 active+clean+snaptrim
52 active+remapped+backfilling
11 active+remapped+backfill_wait+backfill_toofull
io:
client: 155 KiB/s rd, 125 MiB/s wr, 2 op/s rd, 53 op/s wr
recovery: 969 MiB/s, 330 objects/s
***************************************************
If it helps, I've tried capturing 1/5 debug logs from an OSD.
Not sure, but I think this is the way to follow a thread handling one pg
as it decides to rebalance:
[root@ceph7 ceph]# grep 7f2e569e9700 ceph-osd.312.log | less
2020-09-24 14:44:36.844 7f2e569e9700 1 osd.312 pg_epoch: 106808
pg[5.157ds0( v 106803'6043528 (103919'6040524,106803'6043528]
local-lis/les=102671/102672 n=56293 ec=85890/1818 lis/c 102671/10
2671 les/c/f 102672/102672/0 106808/106808/106808)
[148,508,398,457,256,533,137,469,357,306]p148(0) r=-1 lpr=106808
pi=[102671,106808)/1 luod=0'0 crt=106803'6043528 lcod 106801'6043526 active
mbc={} ps=104] start_peering_interval up
[312,424,369,461,546,525,498,169,251,127] ->
[148,508,398,457,256,533,137,469,357,306], acting
[312,424,369,461,546,525,498,169,251,127] -> [148,508,39
8,457,256,533,137,469,357,306], acting_primary 312(0) -> 148, up_primary
312(0) -> 148, role 0 -> -1, features acting 4611087854031667199
upacting 4611087854031667199
2020-09-24 14:44:36.847 7f2e569e9700 1 osd.312 pg_epoch: 106808
pg[5.157ds0( v 106803'6043528 (103919'6040524,106803'6043528]
local-lis/les=102671/102672 n=56293 ec=85890/1818 lis/c 102671/102671
les/c/f 102672/102672/0 106808/106808/106808)
[148,508,398,457,256,533,137,469,357,306]p148(0) r=-1 lpr=106808
pi=[102671,106808)/1 crt=106803'6043528 lcod 106801'6043526 unknown
NOTIFY mbc={} ps=104] state<Start>: transitioning to Stray
2020-09-24 14:44:37.792 7f2e569e9700 1 osd.312 pg_epoch: 106809
pg[5.157ds0( v 106803'6043528 (103919'6040524,106803'6043528]
local-lis/les=102671/102672 n=56293 ec=85890/1818 lis/c 102671/102671
les/c/f 102672/102672/0 106808/106809/106809)
[148,508,398,457,256,533,137,469,357,306]/[312,424,369,461,546,525,498,169,251,127]p312(0)
r=0 lpr=106809 pi=[102671,106809)/1 crt=106803'6043528 lcod
106801'6043526 mlcod 0'0 remapped NOTIFY mbc={} ps=104]
start_peering_interval up [148,508,398,457,256,533,137,469,357,306] ->
[148,508,398,457,256,533,137,469,357,306], acting
[148,508,398,457,256,533,137,469,357,306] ->
[312,424,369,461,546,525,498,169,251,127], acting_primary 148(0) -> 312,
up_primary 148(0) -> 148, role -1 -> 0, features acting
4611087854031667199 upacting 4611087854031667199
2020-09-24 14:44:37.793 7f2e569e9700 1 osd.312 pg_epoch: 106809
pg[5.157ds0( v 106803'6043528 (103919'6040524,106803'6043528]
local-lis/les=102671/102672 n=56293 ec=85890/1818 lis/c 102671/102671
les/c/f 102672/102672/0 106808/106809/106809)
[148,508,398,457,256,533,137,469,357,306]/[312,424,369,461,546,525,498,169,251,127]p312(0)
r=0 lpr=106809 pi=[102671,106809)/1 crt=106803'6043528 lcod
106801'6043526 mlcod 0'0 remapped mbc={} ps=104] state<Start>:
transitioning to Primary
2020-09-24 14:44:38.832 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.137(6) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.861 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.148(0) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.879 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.256(4) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.894 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.306(9) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.902 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.357(8) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.912 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.398(2) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.923 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.457(3) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.931 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.469(7) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.938 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.508(1) from (0'0,0'0] MAX to
106803'6043528
2020-09-24 14:44:38.947 7f2e569e9700 0 log_channel(cluster) log [DBG] :
5.157ds0 starting backfill to osd.533(5) from (0'0,0'0] MAX to
106803'6043528
***************************************************
any advice appreciated,
Jake
--
Dr Jake Grimmett
Head Of Scientific Computing
MRC Laboratory of Molecular Biology
Francis Crick Avenue,
Cambridge CB2 0QH, UK.
_______________________________________________
ceph-users mailing list -- ceph-users@ceph.io
To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Paul, I think you found the answer! When adding 100 new OSDs to the cluster, I increased both pg and pgp from 4096 to 16,384 ********************************** [root@ceph1 ~]# ceph osd pool set ec82pool pg_num 16384 set pool 5 pg_num to 16384 [root@ceph1 ~]# ceph osd pool set ec82pool pgp_num 16384 set pool 5 pgp_num to 16384 ********************************** The pg number increased immediately as seen with "ceph -s" But unknown to me, the pgp number did not increase immediately. "ceph osd pool ls detail" shows that pgp is currently 11412 Each time we hit 5.000% misplaced, the pgp number increases by 1 or 2, this causes the % misplaced to increase again to ~5.1% ... which is why we thought the cluster was not re-balancing. If I'd looked at the ceph.audit.log there are entries like this: 2020-09-23 01:13:11.564384 mon.ceph3b (mon.1) 50747 : audit [INF] from='mgr.90414409 10.1.0.80:0/7898' entity='mgr.ceph2' cmd=[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]: dispatch 2020-09-23 01:13:11.565598 mon.ceph1b (mon.0) 85947 : audit [INF] from='mgr.90414409 ' entity='mgr.ceph2' cmd=[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]: dispatch 2020-09-23 01:13:12.530584 mon.ceph1b (mon.0) 85949 : audit [INF] from='mgr.90414409 ' entity='mgr.ceph2' cmd='[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]': finished Our assumption is that the pgp number will continue to increase till it reaches its set level, at which point the cluster will complete it's re-balance... again, many thanks to you both for your help, Jake On 28/09/2020 17:35, Paul Emmerich wrote:
Hi,
5% misplaced is the default target ratio for misplaced PGs when any automated rebalancing happens, the sources for this are either the balancer or pg scaling. So I'd suspect that there's a PG change ongoing (either pg autoscaler or a manual change, both obey the target misplaced ratio). You can check this by running "ceph osd pool ls detail" and check for the value of pg target.
Also: Looks like you've set osd_scrub_during_recovery = false, this setting can be annoying on large erasure-coded setups on HDDs that see long recovery times. It's better to get IO priorities right; search mailing list for osd op queue cut off high.
Paul
-- Dr Jake Grimmett Head Of Scientific Computing MRC Laboratory of Molecular Biology Francis Crick Avenue, Cambridge CB2 0QH, UK.
On Tue, Sep 29, 2020 at 12:34 PM Jake Grimmett <jog@mrc-lmb.cam.ac.uk> wrote:
I think you found the answer!
When adding 100 new OSDs to the cluster, I increased both pg and pgp from 4096 to 16,384
Too much for your cluster, 4096 seems sufficient for a pool of size 10. You can still reduce it relatively cheaply while it hasn't been fully actuated yet Paul
********************************** [root@ceph1 ~]# ceph osd pool set ec82pool pg_num 16384 set pool 5 pg_num to 16384
[root@ceph1 ~]# ceph osd pool set ec82pool pgp_num 16384 set pool 5 pgp_num to 16384
**********************************
The pg number increased immediately as seen with "ceph -s"
But unknown to me, the pgp number did not increase immediately.
"ceph osd pool ls detail" shows that pgp is currently 11412
Each time we hit 5.000% misplaced, the pgp number increases by 1 or 2, this causes the % misplaced to increase again to ~5.1% ... which is why we thought the cluster was not re-balancing.
If I'd looked at the ceph.audit.log there are entries like this:
2020-09-23 01:13:11.564384 mon.ceph3b (mon.1) 50747 : audit [INF] from='mgr.90414409 10.1.0.80:0/7898' entity='mgr.ceph2' cmd=[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]: dispatch 2020-09-23 01:13:11.565598 mon.ceph1b (mon.0) 85947 : audit [INF] from='mgr.90414409 ' entity='mgr.ceph2' cmd=[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]: dispatch 2020-09-23 01:13:12.530584 mon.ceph1b (mon.0) 85949 : audit [INF] from='mgr.90414409 ' entity='mgr.ceph2' cmd='[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]': finished
Our assumption is that the pgp number will continue to increase till it reaches its set level, at which point the cluster will complete it's re-balance...
again, many thanks to you both for your help,
Jake
On 28/09/2020 17:35, Paul Emmerich wrote:
Hi,
5% misplaced is the default target ratio for misplaced PGs when any automated rebalancing happens, the sources for this are either the balancer or pg scaling. So I'd suspect that there's a PG change ongoing (either pg autoscaler or a manual change, both obey the target misplaced ratio). You can check this by running "ceph osd pool ls detail" and check for the value of pg target.
Also: Looks like you've set osd_scrub_during_recovery = false, this setting can be annoying on large erasure-coded setups on HDDs that see long recovery times. It's better to get IO priorities right; search mailing list for osd op queue cut off high.
Paul
-- Dr Jake Grimmett Head Of Scientific Computing MRC Laboratory of Molecular Biology Francis Crick Avenue, Cambridge CB2 0QH, UK.
I think you found the answer!
When adding 100 new OSDs to the cluster, I increased both pg and pgp from 4096 to 16,384
Too much for your cluster, 4096 seems sufficient for a pool of size 10. You can still reduce it relatively cheaply while it hasn't been fully actuated yet
Assuming he has just the one pool and assuming from the name an 8,2 EC profile, that makes his new PG ratio ~298. Which would not be entirely unreasonable for SSDs, probably too much for HDDs in which case 8192 would arguably be okay. 4096 would make the ratio just 74. — aad
Continuing on this topic, is it only possible to increase the count of placement group (PG) quickly, but the associated placement group placeholder (PGP) values can only increase in smaller increments of 1-3? Each increase of the PGP requires a rebalancing and backfill again of lots of PGs? I am working with an erasure coded pool that I recently increased a target PG count by use of the pg_autoscaler `target_size_ratio` to be closer to what I expect the pool's data size to grow to. I am wondering if this pool will constantly hit 5% misplaced, incrementally add PGP count, and repeat backfilling PGs while the PGs sit unscrubbed. I have 160 OSDs in my pool and now have a target PG count of 2048. I've seen a similar 5% misplaced and never-ending backfills described in a thread by Paul Mezannini (https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/message/OUEZJAEKFV7...). I'd be nice to know the correct strategy to adjust the overall PG count and have a pool return to a healthy + balanced state. -Matt On Tue, Sep 29, 2020 at 9:10 AM Paul Emmerich <emmerich@google.com> wrote:
On Tue, Sep 29, 2020 at 12:34 PM Jake Grimmett <jog@mrc-lmb.cam.ac.uk> wrote:
I think you found the answer!
When adding 100 new OSDs to the cluster, I increased both pg and pgp from 4096 to 16,384
Too much for your cluster, 4096 seems sufficient for a pool of size 10. You can still reduce it relatively cheaply while it hasn't been fully actuated yet
Paul
********************************** [root@ceph1 ~]# ceph osd pool set ec82pool pg_num 16384 set pool 5 pg_num to 16384
[root@ceph1 ~]# ceph osd pool set ec82pool pgp_num 16384 set pool 5 pgp_num to 16384
**********************************
The pg number increased immediately as seen with "ceph -s"
But unknown to me, the pgp number did not increase immediately.
"ceph osd pool ls detail" shows that pgp is currently 11412
Each time we hit 5.000% misplaced, the pgp number increases by 1 or 2, this causes the % misplaced to increase again to ~5.1% ... which is why we thought the cluster was not re-balancing.
If I'd looked at the ceph.audit.log there are entries like this:
2020-09-23 01:13:11.564384 mon.ceph3b (mon.1) 50747 : audit [INF] from='mgr.90414409 10.1.0.80:0/7898' entity='mgr.ceph2' cmd=[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]: dispatch 2020-09-23 01:13:11.565598 mon.ceph1b (mon.0) 85947 : audit [INF] from='mgr.90414409 ' entity='mgr.ceph2' cmd=[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]: dispatch 2020-09-23 01:13:12.530584 mon.ceph1b (mon.0) 85949 : audit [INF] from='mgr.90414409 ' entity='mgr.ceph2' cmd='[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]': finished
Our assumption is that the pgp number will continue to increase till it reaches its set level, at which point the cluster will complete it's re-balance...
again, many thanks to you both for your help,
Jake
On 28/09/2020 17:35, Paul Emmerich wrote:
Hi,
5% misplaced is the default target ratio for misplaced PGs when any automated rebalancing happens, the sources for this are either the balancer or pg scaling. So I'd suspect that there's a PG change ongoing (either pg autoscaler or a manual change, both obey the target misplaced ratio). You can check this by running "ceph osd pool ls detail" and check for the value of pg target.
Also: Looks like you've set osd_scrub_during_recovery = false, this setting can be annoying on large erasure-coded setups on HDDs that see long recovery times. It's better to get IO priorities right; search mailing list for osd op queue cut off high.
Paul
-- Dr Jake Grimmett Head Of Scientific Computing MRC Laboratory of Molecular Biology Francis Crick Avenue, Cambridge CB2 0QH, UK.
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Matt Larson, PhD Madison, WI 53705 U.S.A.
Hi, I’ve just read a post that describe the exact behavior you describe. https://ceph.io/rados/new-in-nautilus-pg-merging-and-autotuning/ There is a config option named target_max_misplaced_ratio, which defaults to 5%. You can change this to accelerate the remap process. Hopes that’s helpful. Sent from my iPad On Sep 29, 2020, at 18:34, Jake Grimmett <jog@mrc-lmb.cam.ac.uk> wrote: Hi Paul, I think you found the answer! When adding 100 new OSDs to the cluster, I increased both pg and pgp from 4096 to 16,384 ********************************** [root@ceph1 ~]# ceph osd pool set ec82pool pg_num 16384 set pool 5 pg_num to 16384 [root@ceph1 ~]# ceph osd pool set ec82pool pgp_num 16384 set pool 5 pgp_num to 16384 ********************************** The pg number increased immediately as seen with "ceph -s" But unknown to me, the pgp number did not increase immediately. "ceph osd pool ls detail" shows that pgp is currently 11412 Each time we hit 5.000% misplaced, the pgp number increases by 1 or 2, this causes the % misplaced to increase again to ~5.1% ... which is why we thought the cluster was not re-balancing. If I'd looked at the ceph.audit.log there are entries like this: 2020-09-23 01:13:11.564384 mon.ceph3b (mon.1) 50747 : audit [INF] from='mgr.90414409 10.1.0.80:0/7898' entity='mgr.ceph2' cmd=[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]: dispatch 2020-09-23 01:13:11.565598 mon.ceph1b (mon.0) 85947 : audit [INF] from='mgr.90414409 ' entity='mgr.ceph2' cmd=[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]: dispatch 2020-09-23 01:13:12.530584 mon.ceph1b (mon.0) 85949 : audit [INF] from='mgr.90414409 ' entity='mgr.ceph2' cmd='[{"prefix": "osd pool set", "pool": "ec82pool", "var": "pgp_num_actual", "val": "5076"}]': finished Our assumption is that the pgp number will continue to increase till it reaches its set level, at which point the cluster will complete it's re-balance... again, many thanks to you both for your help, Jake On 28/09/2020 17:35, Paul Emmerich wrote: Hi, 5% misplaced is the default target ratio for misplaced PGs when any automated rebalancing happens, the sources for this are either the balancer or pg scaling. So I'd suspect that there's a PG change ongoing (either pg autoscaler or a manual change, both obey the target misplaced ratio). You can check this by running "ceph osd pool ls detail" and check for the value of pg target. Also: Looks like you've set osd_scrub_during_recovery = false, this setting can be annoying on large erasure-coded setups on HDDs that see long recovery times. It's better to get IO priorities right; search mailing list for osd op queue cut off high. Paul -- Dr Jake Grimmett Head Of Scientific Computing MRC Laboratory of Molecular Biology Francis Crick Avenue, Cambridge CB2 0QH, UK. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Dear All, great advice - thank you all so much. I've changed the pgp to 8192 (it had already risen to 11857) and see how this works. The target_max_misplaced_ratio looks like a useful control. It's a shame the ceph pg calc page <https://ceph.io/pgcalc/> doesn't have more advice for people using erasure pools... We are planning to go from 550 osd to 700 osds soon, and eventually to 900 osds. What is the ideal pg count for 900 osds on an EC 8+2 pool that will probably reach 80% full? best regards, Jake On 30/09/2020 04:50, 胡 玮文 wrote:
Hi,
I’ve just read a post that describe the exact behavior you describe. https://ceph.io/rados/new-in-nautilus-pg-merging-and-autotuning/
There is a config option named /target_max_misplaced_ratio/, which defaults to 5%. You can change this to accelerate the remap process.
Hopes that’s helpful.
Sent from my iPad
On Sep 29, 2020, at 18:34, Jake Grimmett <jog@mrc-lmb.cam.ac.uk> wrote:
Hi Paul,
I think you found the answer!
When adding 100 new OSDs to the cluster, I increased both pg and pgp from 4096 to 16,384
********************************** [root@ceph1 ~]# ceph osd pool set ec82pool pg_num 16384 set pool 5 pg_num to 16384
[root@ceph1 ~]# ceph osd pool set ec82pool pgp_num 16384 set pool 5 pgp_num to 16384
**********************************
The pg number increased immediately as seen with "ceph -s"
But unknown to me, the pgp number did not increase immediately.
"ceph osd pool ls detail" shows that pgp is currently 11412
Each time we hit 5.000% misplaced, the pgp number increases by 1 or 2, this causes the % misplaced to increase again to ~5.1% -- Dr Jake Grimmett Head Of Scientific Computing MRC Laboratory of Molecular Biology Francis Crick Avenue, Cambridge CB2 0QH, UK.
Does changing `target_max_misplaced_ratio` result in more PGP being created in each cycle of the remapping? Would this result in fewer copies of data or just more PGs being processed in each batch during a change of PG numbers? What is a safe value to raise `target_max_misplaced_ratio` to, given that the default is 0.05 or 5% misplaced objects? -Matt On Wed, Sep 30, 2020 at 3:20 AM Jake Grimmett <jog@mrc-lmb.cam.ac.uk> wrote:
Dear All,
great advice - thank you all so much.
I've changed the pgp to 8192 (it had already risen to 11857) and see how this works. The target_max_misplaced_ratio looks like a useful control.
It's a shame the ceph pg calc page <https://ceph.io/pgcalc/> doesn't have more advice for people using erasure pools...
We are planning to go from 550 osd to 700 osds soon, and eventually to 900 osds.
What is the ideal pg count for 900 osds on an EC 8+2 pool that will probably reach 80% full?
best regards,
Jake
On 30/09/2020 04:50, 胡 玮文 wrote:
Hi,
I’ve just read a post that describe the exact behavior you describe. https://ceph.io/rados/new-in-nautilus-pg-merging-and-autotuning/
There is a config option named /target_max_misplaced_ratio/, which defaults to 5%. You can change this to accelerate the remap process.
Hopes that’s helpful.
Sent from my iPad
On Sep 29, 2020, at 18:34, Jake Grimmett <jog@mrc-lmb.cam.ac.uk> wrote:
Hi Paul,
I think you found the answer!
When adding 100 new OSDs to the cluster, I increased both pg and pgp from 4096 to 16,384
********************************** [root@ceph1 ~]# ceph osd pool set ec82pool pg_num 16384 set pool 5 pg_num to 16384
[root@ceph1 ~]# ceph osd pool set ec82pool pgp_num 16384 set pool 5 pgp_num to 16384
**********************************
The pg number increased immediately as seen with "ceph -s"
But unknown to me, the pgp number did not increase immediately.
"ceph osd pool ls detail" shows that pgp is currently 11412
Each time we hit 5.000% misplaced, the pgp number increases by 1 or 2, this causes the % misplaced to increase again to ~5.1% -- Dr Jake Grimmett Head Of Scientific Computing MRC Laboratory of Molecular Biology Francis Crick Avenue, Cambridge CB2 0QH, UK.
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Matt Larson, PhD Madison, WI 53705 U.S.A.
participants (6)
-
Anthony D'Atri
-
Jake Grimmett
-
Matt Larson
-
Paul Emmerich
-
Stefan Kooman
-
胡 玮文