Hi I tried to enable the new ec optimisation flag on one of the erasure encoded pools for rbd on my recently upgraded to Tentacle 20.2.0 version ceph cluster. ceph osd pool set rbd_ecpool allow_ec_optimizations true At first everything seemed normal. After a while when I went back to the computer I saw that osd.1 and osd.3 were always crashing. After a reboot I then saw errors and warnings about inconsistent pg's. OSD_SCRUB_ERRORS: 116116 scrub errors OSD_TOO_MANY_REPAIRS: Too many repaired reads on 4 OSDs PG_DAMAGED: Possible data damage: 3 pgs inconsistent I tried to run a manual ceph pg repair <pgid> on the affected pg's. While the affected pg's disappeared from the inconsistent pg list after a deep-scrub+repair new pgid appeared on the incosistent pg list and the amount of scrub errors increased. When I run a manual repair on a pgid, I see a lot of errors in the pg's primary osd log like: - 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39(237963'11927123 client.40141308.0:1785630 dirty s 4194304 uv 11910481 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : size 1400832 != size 1396736 from auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39(237963'11927123 client.40141308.0:1785630 dirty s 4194304 uv 11910481 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45(238043'11948425 client.40186030.0:433080 dirty s 4194304 uv 11933347 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : size 1400832 != size 1396736 from auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45(238043'11948425 client.40186030.0:433080 dirty s 4194304 uv 11933347 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d(238319'11972769 client.40186030.0:1294524 dirty|data_digest s 4194304 uv 11955052 dd 2cef26e9 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : size 1400832 != size 1396736 The size mismatch always seem to be the same: 1400832 != size 1396736 And looking at the rbd namespace like example rbd_data.7.6f5c183c7dc42b It looks like so far that images that have been in use/mounted during enablement of allow_ec_pool_optimization are affected. So my questions are: 1. Are these errors expected/explainable? 2. Will a manual pg repair on all the pools pgids eventually fix the problem? 3. Is there any chance that data has been lost/inconsistent after pg repair on all the pools pgids? Thanks & best regards, Reto