Hi, The cluster is with Quincy, managed by cephadm, 70 nodes, 1600 SSD OSDs, 5 mons. When creating an EC 4+2 pool, log says crush smoke test takes 123s. 2025-11-27T05:10:17.006+0000 7f335487f700 10 mon.ceph-1@0(leader).osd e3155117 prepare_new_pool crush test_with_fork tester created 2025-11-27T05:10:17.006+0000 7f335487f700 10 mon.ceph-1@0(leader).osd e3155117 prepare_new_pool crush smoke test duration: 123.727737427s The pool seems being created. 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 scan_for_creating_pgs queueing pool create for 73 erasure profile nvme-ec42 size 6 min_size 5 crush_rule 4 object_hash rjenkins pg_num 1 pgp_num 1 autoscale_mode on last_change 3155118 flags hashpspool,creating stripe_width 16384 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 update_pending_pgs 1 pools queued 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 update_pending_pgs 0 pgs removed because they're created 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 update_pending_pgs pool 73 created 3155118 modified 2025-11-27T05:10:17.010928+0000 [0-1) 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 update_pending_pgs adding 73.0 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 update_pending_pgs done with queue for 73 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 update_pending_pgs queue remaining: 0 pools 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 update_pending_pgs pg 73.0 just added, up [862,1320,166,422,263,984] p 862 acting [862,1320,166,422,263,984] p 862 history ec=3155118/3155118 lis/c=0/0 les/c/f=0/0/0 sis=31551 18 past_intervals ([0,0] all_participants= intervals=) 2025-11-27T05:10:17.114+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 update_pending_pgs 1/1 pgs added from queued pools 2025-11-27T05:10:17.115+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 encode_pending encoding full map with quincy features 1080873256688364036 2025-11-27T05:10:17.117+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 insert_purged_snap_update [1921e6,1921e7) - join with earlier [1921dc,1921e6) 2025-11-27T05:10:17.117+0000 7f3357084700 10 mon.ceph-1@0(leader).osd e3155117 insert_purged_snap_update [1a18b8,1a18b9) - join with earlier [1a18b7,1a18b8) 2025-11-27T05:10:17.117+0000 7f3357084700 10 mon.ceph-1@0(leader) e17 log_health updated 0 previous 0 Then there is complaints about this slowness. 2025-11-27T05:10:17.123+0000 7f3357084700 1 mon.ceph-1@0(leader).mds e1 check_health: resetting beacon timeouts due to mon delay (slow election?) of 126.641 seconds 2025-11-27T05:10:17.123+0000 7f3357084700 10 mon.ceph-1@0(leader).log v37135378 log 2025-11-27T05:10:17.123+0000 7f3357084700 10 mon.ceph-1@0(leader).auth v111171 auth 2025-11-27T05:10:17.123+0000 7f3357084700 4 mon.ceph-1@0(leader).mgr e663 tick: resetting beacon timeouts due to mon delay (slow election?) of 126.532714844s seconds 2025-11-27T05:10:17.124+0000 7f3357084700 10 mon.ceph-1@0(leader).health tick 2025-11-27T05:10:17.124+0000 7f3357084700 10 mon.ceph-1@0(leader).health check_member_health avail 91% total 444 GiB, used 38 GiB, avail 407 GiB 2025-11-27T05:10:17.125+0000 7f3357084700 10 mon.ceph-1@0(leader).health 7971095 2025-11-27T05:10:17.125+0000 7f3357084700 10 mon.ceph-1@0(leader) e17 log_health updated 0 previous 0 2025-11-27T05:10:17.125+0000 7f3357084700 10 mon.ceph-1@0(leader).config tick 2025-11-27T05:10:17.125+0000 7f3357084700 10 mon.ceph-1@0(leader).kv tick 2025-11-27T05:10:17.125+0000 7f3357084700 -1 mon.ceph-1@0(leader) e17 get_health_metrics reporting 3 slow ops, oldest is pool_op(delete unmanaged snap pool 6 tid 22682908 name v3154700) Then it triggered mon re-election. From CLI, the "ceph osd pool create" command is stuck forever. Mon.1 keeps being out and rejoining the cluster. Eventually, Ctrl-C breaks the CLI command, mon.1 stops in-and-out, but the pool is not created. Is it normal to take so long for smoke test when creating EC pool in such cluster size? Any way to make the creation faster? Or Any way to increase timeout to avoid re-election in such case? Thanks! Tony