[Tentacle 20.2.0]: Inconsistent pg's after enabling ec optimisation flag
Hi I tried to enable the new ec optimisation flag on one of the erasure encoded pools for rbd on my recently upgraded to Tentacle 20.2.0 version ceph cluster. ceph osd pool set rbd_ecpool allow_ec_optimizations true At first everything seemed normal. After a while when I went back to the computer I saw that osd.1 and osd.3 were always crashing. After a reboot I then saw errors and warnings about inconsistent pg's. OSD_SCRUB_ERRORS: 116116 scrub errors OSD_TOO_MANY_REPAIRS: Too many repaired reads on 4 OSDs PG_DAMAGED: Possible data damage: 3 pgs inconsistent I tried to run a manual ceph pg repair <pgid> on the affected pg's. While the affected pg's disappeared from the inconsistent pg list after a deep-scrub+repair new pgid appeared on the incosistent pg list and the amount of scrub errors increased. When I run a manual repair on a pgid, I see a lot of errors in the pg's primary osd log like: - 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39(237963'11927123 client.40141308.0:1785630 dirty s 4194304 uv 11910481 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : size 1400832 != size 1396736 from auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39(237963'11927123 client.40141308.0:1785630 dirty s 4194304 uv 11910481 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45(238043'11948425 client.40186030.0:433080 dirty s 4194304 uv 11933347 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : size 1400832 != size 1396736 from auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45(238043'11948425 client.40186030.0:433080 dirty s 4194304 uv 11933347 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d(238319'11972769 client.40186030.0:1294524 dirty|data_digest s 4194304 uv 11955052 dd 2cef26e9 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : size 1400832 != size 1396736 The size mismatch always seem to be the same: 1400832 != size 1396736 And looking at the rbd namespace like example rbd_data.7.6f5c183c7dc42b It looks like so far that images that have been in use/mounted during enablement of allow_ec_pool_optimization are affected. So my questions are: 1. Are these errors expected/explainable? 2. Will a manual pg repair on all the pools pgids eventually fix the problem? 3. Is there any chance that data has been lost/inconsistent after pg repair on all the pools pgids? Thanks & best regards, Reto
Update: Now some OSD's are back to crashing again and again. Here's one of the crash dumps: -10> 2025-11-29T15:46:00.712+0000 7fb77f9ee640 5 osd.1 pg_epoch: 240317 pg[90.bs0( v 240049'12124879 (239808'12122484,240049'12124879] local-lis/les=240316/240317 n=20317 ec=74403/72901 lis/c=240316/240292 les/c/f=240317/240293/0 sis=240316) [1,15,3,9,0]p1(0) r=0 lpr=240316 pi=[240292,240316)/1 crt=240049'12124 879 lcod 0'0 mlcod 0'0 active mbc={}] exit Started/Primary/Active/Activating 0.002957 11 0.000191 -9> 2025-11-29T15:46:00.712+0000 7fb77f9ee640 5 osd.1 pg_epoch: 240317 pg[90.bs0( v 240049'12124879 (239808'12122484,240049'12124879] local-lis/les=240316/240317 n=20317 ec=74403/72901 lis/c=240316/240292 les/c/f=240317/240293/0 sis=240316) [1,15,3,9,0]p1(0) r=0 lpr=240316 pi=[240292,240316)/1 crt=240049'12124 879 lcod 0'0 mlcod 0'0 active mbc={}] enter Started/Primary/Active/Recovered -8> 2025-11-29T15:46:00.712+0000 7fb77f9ee640 5 osd.1 pg_epoch: 240317 pg[90.bs0( v 240049'12124879 (239808'12122484,240049'12124879] local-lis/les=240316/240317 n=20317 ec=74403/72901 lis/c=240316/240292 les/c/f=240317/240293/0 sis=240316) [1,15,3,9,0]p1(0) r=0 lpr=240316 pi=[240292,240316)/1 crt=240049'12124 879 lcod 0'0 mlcod 0'0 active mbc={}] exit Started/Primary/Active/Recovered 0.000012 0 0.000000 -7> 2025-11-29T15:46:00.712+0000 7fb77f9ee640 5 osd.1 pg_epoch: 240317 pg[90.bs0( v 240049'12124879 (239808'12122484,240049'12124879] local-lis/les=240316/240317 n=20317 ec=74403/72901 lis/c=240316/240292 les/c/f=240317/240293/0 sis=240316) [1,15,3,9,0]p1(0) r=0 lpr=240316 pi=[240292,240316)/1 crt=240049'12124 879 lcod 0'0 mlcod 0'0 active mbc={}] enter Started/Primary/Active/Clean -6> 2025-11-29T15:46:00.712+0000 7fb7809f0640 1 osd.1 pg_epoch: 240317 pg[90.12s3( v 240041'11447066 (239802'11445030,240041'11447066] local-lis/les=240316/240317 n=20720 ec=183839/72901 lis/c=240307/240132 les/c/f=240308/240133/0 sis=240316) [NONE,8,9,1,0]p1(3) r=3 lpr=240316 pi=[240132,240316)/2 crt=240041'1 1447066 lcod 0'0 mlcod 0'0 active+undersized+degraded mbc={}] state<Started/Primary/Active>: react AllReplicasActivated Activating complete -5> 2025-11-29T15:46:00.712+0000 7fb7809f0640 5 osd.1 pg_epoch: 240317 pg[90.12s3( v 240041'11447066 (239802'11445030,240041'11447066] local-lis/les=240316/240317 n=20720 ec=183839/72901 lis/c=240316/240132 les/c/f=240317/240133/0 sis=240316) [NONE,8,9,1,0]p1(3) r=3 lpr=240316 pi=[240132,240316)/2 crt=240041'1 1447066 lcod 0'0 mlcod 0'0 active+undersized+degraded mbc={}] exit Started/Primary/Active/Activating 0.002216 9 0.000135 -4> 2025-11-29T15:46:00.712+0000 7fb7809f0640 5 osd.1 pg_epoch: 240317 pg[90.12s3( v 240041'11447066 (239802'11445030,240041'11447066] local-lis/les=240316/240317 n=20720 ec=183839/72901 lis/c=240316/240132 les/c/f=240317/240133/0 sis=240316) [NONE,8,9,1,0]p1(3) r=3 lpr=240316 pi=[240132,240316)/2 crt=240041'1 1447066 lcod 0'0 mlcod 0'0 active+undersized+degraded mbc={}] enter Started/Primary/Active/Recovered -3> 2025-11-29T15:46:00.712+0000 7fb7809f0640 5 osd.1 pg_epoch: 240317 pg[90.12s3( v 240041'11447066 (239802'11445030,240041'11447066] local-lis/les=240316/240317 n=20720 ec=183839/72901 lis/c=240316/240132 les/c/f=240317/240133/0 sis=240316) [NONE,8,9,1,0]p1(3) r=3 lpr=240316 pi=[240132,240316)/2 crt=240041'1 1447066 lcod 0'0 mlcod 0'0 active+undersized+degraded mbc={}] exit Started/Primary/Active/Recovered 0.000004 0 0.000000 -2> 2025-11-29T15:46:00.712+0000 7fb7809f0640 5 osd.1 pg_epoch: 240317 pg[90.12s3( v 240041'11447066 (239802'11445030,240041'11447066] local-lis/les=240316/240317 n=20720 ec=183839/72901 lis/c=240316/240132 les/c/f=240317/240133/0 sis=240316) [NONE,8,9,1,0]p1(3) r=3 lpr=240316 pi=[240132,240316)/2 crt=240041'1 1447066 lcod 0'0 mlcod 0'0 active+undersized+degraded mbc={}] enter Started/Primary/Active/Clean -1> 2025-11-29T15:46:00.716+0000 7fb78bdc7640 3 osd.1 240317 handle_osd_map epochs [240317,240317], i have 240317, src has [239665,240317] 0> 2025-11-29T15:46:00.716+0000 7fb7801ef640 -1 *** Caught signal (Aborted) ** in thread 7fb7801ef640 thread_name:tp_osd_tp ceph version 20.2.0 (69f84cc2651aa259a15bc192ddaabd3baba07489) tentacle (stable - RelWithDebInfo) 1: /lib64/libc.so.6(+0x3fc30) [0x7fb79c33dc30] 2: /lib64/libc.so.6(+0x8d03c) [0x7fb79c38b03c] 3: raise() 4: abort() 5: /lib64/libstdc++.so.6(+0xa1b21) [0x7fb79d0dab21] 6: /lib64/libstdc++.so.6(+0xad53c) [0x7fb79d0e653c] 7: /lib64/libstdc++.so.6(+0xad5a7) [0x7fb79d0e65a7] 8: /lib64/libstdc++.so.6(+0xad809) [0x7fb79d0e6809] 9: /usr/bin/ceph-osd(+0x41b012) [0x564a726e4012] 10: (ECTransaction::WritePlanObj::WritePlanObj(hobject_t const&, PGTransaction::ObjectOperation const&, ECUtil::stripe_info_t const&, bitset_set<128ul, shard_id_t>, bitset_set<128ul, shard_id_t>, bool, unsigned long, std::optional<object_info_t> const&, std::optional<object_info_t> const&, unsigned int)+0x1314) [0 x564a72c841c4] 11: /usr/bin/ceph-osd(+0x99cdaa) [0x564a72c65daa] 12: (ECCommon::get_write_plan(ECUtil::stripe_info_t const&, PGTransaction&, ECCommon::ReadPipeline&, ECCommon::RMWPipeline&, DoutPrefixProvider*)+0x54f) [0x564a72c669ff] 13: (ECBackend::submit_transaction(hobject_t const&, object_stat_sum_t const&, eversion_t const&, std::unique_ptr<PGTransaction, std::default_delete<PGTransaction> >&&, eversion_t const&, eversion_t const&, std::vector<pg_log_entry_t, std::allocator<pg_log_entry_t> >&&, std::optional<pg_hit_set_history_t>&, Contex t*, unsigned long, osd_reqid_t, boost::intrusive_ptr<OpRequest>)+0x62e) [0x564a72c7b79e] 14: /usr/bin/ceph-osd(+0x7d980f) [0x564a72aa280f] 15: (PrimaryLogPG::issue_repop(PrimaryLogPG::RepGather*, PrimaryLogPG::OpContext*)+0x3ae) [0x564a72a2302e] 16: (PrimaryLogPG::execute_ctx(PrimaryLogPG::OpContext*)+0xf6a) [0x564a72a0063a] 17: (PrimaryLogPG::do_op(boost::intrusive_ptr<OpRequest>&)+0x2d5f) [0x564a729f0f5f] 18: (OSD::dequeue_op(boost::intrusive_ptr<PG>, boost::intrusive_ptr<OpRequest>, ThreadPool::TPHandle&)+0x19f) [0x564a7292005f] 19: (ceph::osd::scheduler::PGOpItem::run(OSD*, OSDShard*, boost::intrusive_ptr<PG>&, ThreadPool::TPHandle&)+0x69) [0x564a72b80f89] 20: (OSD::ShardedOpWQ::_process(unsigned int, ceph::heartbeat_handle_d*)+0x8bc) [0x564a7294911c] 21: (ShardedThreadPool::shardedthreadpool_worker(unsigned int)+0x23a) [0x564a72ecbe1a] 22: /usr/bin/ceph-osd(+0xc033d4) [0x564a72ecc3d4] 23: /lib64/libc.so.6(+0x8b2fa) [0x7fb79c3892fa] 24: /lib64/libc.so.6(+0x110400) [0x7fb79c40e400] NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this. --- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 3/ 5 mds_quiesce 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 rgw_lifecycle 1/ 5 rgw_restore 1/ 5 rgw_notification 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluestore_compression 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 crimson_interrupt 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 1/ 5 rgw_dedup 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore 1/ 5 trace 0/ 5 ceph_dedup -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7fb77e1eb640 / osd_srv_heartbt 7fb77e9ec640 / tp_osd_tp 7fb77f1ed640 / tp_osd_tp 7fb77f9ee640 / tp_osd_tp 7fb7801ef640 / tp_osd_tp 7fb7809f0640 / tp_osd_tp 7fb789a02640 / ceph-osd 7fb78a203640 / cfin 7fb78aa04640 / bstore_kv_sync 7fb78bdc7640 / ms_dispatch 7fb78cdc9640 / safe_timer 7fb78ddcb640 / ms_dispatch 7fb78ea0c640 / bstore_mempool 7fb78fc21640 / rocksdb:low 7fb790c23640 / rocksdb:low 7fb792c27640 / fn_anonymous 7fb79442a640 / safe_timer 7fb797e11640 / io_context_pool 7fb798656640 / io_context_pool 7fb799658640 / admin_socket, ms_dispatch, admin_socket, ms_dispatch 7fb799e59640 / msgr-worker-2 7fb79a65a640 / msgr-worker-1 7fb79ae5b640 / msgr-worker-0 7fb79bec28c0 / ceph-osd max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-osd.1.log --- end dump of recent events --- Am Sa., 29. Nov. 2025 um 13:11 Uhr schrieb Reto Gysi <rlgysi@gmail.com>:
Hi
I tried to enable the new ec optimisation flag on one of the erasure encoded pools for rbd on my recently upgraded to Tentacle 20.2.0 version ceph cluster.
ceph osd pool set rbd_ecpool allow_ec_optimizations true
At first everything seemed normal. After a while when I went back to the computer I saw that osd.1 and osd.3 were always crashing. After a reboot I then saw errors and warnings about inconsistent pg's.
OSD_SCRUB_ERRORS: 116116 scrub errors OSD_TOO_MANY_REPAIRS: Too many repaired reads on 4 OSDs PG_DAMAGED: Possible data damage: 3 pgs inconsistent
I tried to run a manual ceph pg repair <pgid> on the affected pg's. While the affected pg's disappeared from the inconsistent pg list after a deep-scrub+repair new pgid appeared on the incosistent pg list and the amount of scrub errors increased.
When I run a manual repair on a pgid, I see a lot of errors in the pg's primary osd log like:
- 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39(237963'11927123 client.40141308.0:1785630 dirty s 4194304 uv 11910481 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39 : size 1400832 != size 1396736 from auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d39(237963'11927123 client.40141308.0:1785630 dirty s 4194304 uv 11910481 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45(238043'11948425 client.40186030.0:433080 dirty s 4194304 uv 11933347 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45 : size 1400832 != size 1396736 from auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d45(238043'11948425 client.40186030.0:433080 dirty s 4194304 uv 11933347 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : candidate size 1400832 info size 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.b shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : candidate size 1400832 info siz e 1396736 mismatch 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 3(2) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : size 1400832 != size 1396736 f rom auth oi 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d(238319'11972769 client.40186030.0:1294524 dirty|data_digest s 4194304 uv 11955052 dd 2cef26e9 alloc_hint [0 0 0]) 2025-11-29T11:48:10.251+0000 7f7328665640 -1 log_channel(cluster) log [ERR] : 90.bs0 shard 15(1) soid 90:d0bc306e:::rbd_data.7.6f5c183c7dc42b.0000000000001331:2d5d : size 1400832 != size 1396736
The size mismatch always seem to be the same: 1400832 != size 1396736
And looking at the rbd namespace like example rbd_data.7.6f5c183c7dc42b It looks like so far that images that have been in use/mounted during enablement of allow_ec_pool_optimization are affected.
So my questions are: 1. Are these errors expected/explainable? 2. Will a manual pg repair on all the pools pgids eventually fix the problem? 3. Is there any chance that data has been lost/inconsistent after pg repair on all the pools pgids?
Thanks & best regards,
Reto
participants (1)
-
Reto Gysi