Urgent help with degraded filesystem needed
Hello cephers, we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again. We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby) It started this night, I got the first HEALTH_WARN emails saying: HEALTH_WARN --- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074 === Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074 then it went on with: HEALTH_WARN --- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded --- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074 === Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded Then one after another MDS was going to error state: HEALTH_WARN --- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state === Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1 ceph fs set cephfs max_mds 1 And set the filesystem down ceph fs set cephfs down true I tried to restart the MDS again but now I'm stuck at the following status: [root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M) data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep [root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) The end log file of the replay(laggy) default.cephmon-02.duujba shows: [...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: 568: FAILED ceph_assert(p->first <= start) ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone() 0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this. --- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events --- I have no idea how to resolve this and would be grateful for any help. Dietmar
Hi Dietmar, On 6/19/24 15:43, Dietmar Rieder wrote:
Hello cephers,
we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again.
We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby)
It started this night, I got the first HEALTH_WARN emails saying:
HEALTH_WARN
--- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
then it went on with:
HEALTH_WARN
--- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
--- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
Then one after another MDS was going to error state:
HEALTH_WARN
--- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state
=== Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more
In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1
ceph fs set cephfs max_mds 1
And set the filesystem down
ceph fs set cephfs down true
I tried to restart the MDS again but now I'm stuck at the following status:
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M)
data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep
[root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The end log file of the replay(laggy) default.cephmon-02.duujba shows:
[...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: 568: FAILED ceph_assert(p->first <= start)
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone()
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
This is a known bug, please see https://tracker.ceph.com/issues/61009. As a workaround I am afraid you need to trim the journal logs first and then try to restart the MDS daemons, And at the same time please follow the workaround in https://tracker.ceph.com/issues/61009#note-26 Thanks - Xiubo
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events ---
I have no idea how to resolve this and would be grateful for any help.
Dietmar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Xiubo, On 6/19/24 09:55, Xiubo Li wrote:
Hi Dietmar,
On 6/19/24 15:43, Dietmar Rieder wrote:
Hello cephers,
we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again.
We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby)
It started this night, I got the first HEALTH_WARN emails saying:
HEALTH_WARN
--- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
then it went on with:
HEALTH_WARN
--- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
--- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
Then one after another MDS was going to error state:
HEALTH_WARN
--- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state
=== Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more
In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1
ceph fs set cephfs max_mds 1
And set the filesystem down
ceph fs set cephfs down true
I tried to restart the MDS again but now I'm stuck at the following status:
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M)
data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep
[root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The end log file of the replay(laggy) default.cephmon-02.duujba shows:
[...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: 568: FAILED ceph_assert(p->first <= start)
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone()
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
This is a known bug, please see https://tracker.ceph.com/issues/61009.
As a workaround I am afraid you need to trim the journal logs first and then try to restart the MDS daemons, And at the same time please follow the workaround in https://tracker.ceph.com/issues/61009#note-26
I see, I'll try to do this. Are there any caveats or issues to expect by trimming the journal logs? Is there a step by step guide on how to perform the trimming? Should all MDS be stopped before? Sorry for the lot of (naive) questions, but I do not want to make any mistake here. Thanks for your support, Dietmar
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events ---
I have no idea how to resolve this and would be grateful for any help.
Dietmar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On 6/19/24 16:13, Dietmar Rieder wrote:
Hi Xiubo,
[...]
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
This is a known bug, please see https://tracker.ceph.com/issues/61009.
As a workaround I am afraid you need to trim the journal logs first and then try to restart the MDS daemons, And at the same time please follow the workaround in https://tracker.ceph.com/issues/61009#note-26
I see, I'll try to do this. Are there any caveats or issues to expect by trimming the journal logs?
Certainly you will lose the dirty metadata in the journals.
Is there a step by step guide on how to perform the trimming? Should all MDS be stopped before?
Please follow https://docs.ceph.com/en/nautilus/cephfs/disaster-recovery-experts/#disaster....
Sorry for the lot of (naive) questions, but I do not want to make any mistake here.
Since the journal logs were corrupted and couldn't be replayed by the MDS when starting and the MDS crash will continue unless you manually repair or truncate it. Thanks - Xiubo
Thanks for your support,
Dietmar
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events ---
I have no idea how to resolve this and would be grateful for any help.
Dietmar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On 6/19/24 10:30, Xiubo Li wrote:
On 6/19/24 16:13, Dietmar Rieder wrote:
Hi Xiubo,
[...]
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
This is a known bug, please see https://tracker.ceph.com/issues/61009.
As a workaround I am afraid you need to trim the journal logs first and then try to restart the MDS daemons, And at the same time please follow the workaround in https://tracker.ceph.com/issues/61009#note-26
I see, I'll try to do this. Are there any caveats or issues to expect by trimming the journal logs?
Certainly you will lose the dirty metadata in the journals.
Is there a step by step guide on how to perform the trimming? Should all MDS be stopped before?
Please follow https://docs.ceph.com/en/nautilus/cephfs/disaster-recovery-experts/#disaster....
OK, when I run the cephfs-journal-tool I get an error: # cephfs-journal-tool journal export backup.bin Error ((22) Invalid argument) My cluster is managed by caphadm, so (in my stress situation) I'm not able find the correct way to use cephfs-journal-tool I'm sure it is something stupid that I'm missing but I'd be happy for any hint. Thanks Dietmar
Sorry for the lot of (naive) questions, but I do not want to make any mistake here.
Since the journal logs were corrupted and couldn't be replayed by the MDS when starting and the MDS crash will continue unless you manually repair or truncate it.
Thanks
- Xiubo
Thanks for your support,
Dietmar
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events ---
I have no idea how to resolve this and would be grateful for any help.
Dietmar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi, On 19-06-2024 11:15, Dietmar Rieder wrote:
Please follow https://docs.ceph.com/en/nautilus/cephfs/disaster-recovery-experts/#disaster....
OK, when I run the cephfs-journal-tool I get an error:
# cephfs-journal-tool journal export backup.bin Error ((22) Invalid argument)
My cluster is managed by caphadm, so (in my stress situation) I'm not able find the correct way to use cephfs-journal-tool > I'm sure it is something stupid that I'm missing but I'd be happy for any hint.
cephadm shell --mount /root/ -- cephfs-journal-tool journal export /mnt/backup.bin ^^ Does this work? The backup.bin file should end up in the home dir of the root user. Gr. Stefan
Hi, On 6/19/24 12:14, Stefan Kooman wrote:
Hi,
On 19-06-2024 11:15, Dietmar Rieder wrote:
Please follow https://docs.ceph.com/en/nautilus/cephfs/disaster-recovery-experts/#disaster....
OK, when I run the cephfs-journal-tool I get an error:
# cephfs-journal-tool journal export backup.bin Error ((22) Invalid argument)
My cluster is managed by caphadm, so (in my stress situation) I'm not able find the correct way to use cephfs-journal-tool > I'm sure it is something stupid that I'm missing but I'd be happy for any hint.
cephadm shell --mount /root/ -- cephfs-journal-tool journal export /mnt/backup.bin
^^ Does this work? The backup.bin file should end up in the home dir of the root user.
Thaks, this worked: cephfs-journal-tool --rank=cephfs:all journal export /mnt/backup/backup.bin Dietmar
On 6/19/24 11:15, Dietmar Rieder wrote:
On 6/19/24 10:30, Xiubo Li wrote:
On 6/19/24 16:13, Dietmar Rieder wrote:
Hi Xiubo,
[...]
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
This is a known bug, please see https://tracker.ceph.com/issues/61009.
As a workaround I am afraid you need to trim the journal logs first and then try to restart the MDS daemons, And at the same time please follow the workaround in https://tracker.ceph.com/issues/61009#note-26
I see, I'll try to do this. Are there any caveats or issues to expect by trimming the journal logs?
Certainly you will lose the dirty metadata in the journals.
Is there a step by step guide on how to perform the trimming? Should all MDS be stopped before?
Please follow https://docs.ceph.com/en/nautilus/cephfs/disaster-recovery-experts/#disaster....
OK, when I run the cephfs-journal-tool I get an error:
# cephfs-journal-tool journal export backup.bin Error ((22) Invalid argument)
My cluster is managed by caphadm, so (in my stress situation) I'm not able find the correct way to use cephfs-journal-tool
I'm sure it is something stupid that I'm missing but I'd be happy for any hint.
I ran the disaster recovery procedures now, as follows: [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 event recover_dentries summary Events by type: OPEN: 8737 PURGED: 1 SESSION: 9 SESSIONS: 2 SUBTREEMAP: 128 TABLECLIENT: 2 TABLESERVER: 30 UPDATE: 9207 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 event recover_dentries summary Events by type: OPEN: 3 SESSION: 1 SUBTREEMAP: 34 UPDATE: 32965 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 event recover_dentries summary Events by type: OPEN: 5289 SESSION: 10 SESSIONS: 3 SUBTREEMAP: 128 UPDATE: 76448 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:all journal inspect Overall journal integrity: OK Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 journal inspect Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 journal inspect Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 journal reset old journal was 879331755046~508520587 new journal start will be 879843344384 (3068751 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal reset old journal was 934711229813~120432327 new journal start will be 934834864128 (3201988 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 journal reset old journal was 1334153584288~252692691 new journal start will be 1334409428992 (3152013 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-table-tool all reset session { "0": { "data": {}, "result": 0 }, "1": { "data": {}, "result": 0 }, "2": { "data": {}, "result": 0 } } [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: OK [root@ceph01-b /]# ceph fs reset cephfs --yes-i-really-mean-it But now I hit the error below: -20> 2024-06-19T11:13:00.610+0000 7ff3694d0700 10 monclient: _send_mon_message to mon.cephmon-03 at v2:10.1.3.23:3300/0 -19> 2024-06-19T11:13:00.637+0000 7ff3664ca700 2 mds.0.cache Memory usage: total 485928, rss 170860, heap 207156, baseline 182580, 0 / 33434 inodes have caps, 0 caps, 0 caps per inode -18> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.default.cephmon-03.chjusj Updating MDS map to version 8061 from mon.1 -17> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 handle_mds_map i am now mds.0.8058 -16> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 handle_mds_map state change up:rejoin --> up:active -15> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 recovery_done -- successful recovery! -14> 2024-06-19T11:13:00.788+0000 7ff36a4d2700 1 mds.0.8058 active_start -13> 2024-06-19T11:13:00.789+0000 7ff36dcd9700 5 mds.beacon.default.cephmon-03.chjusj received beacon reply up:active seq 4 rtt 0.955007 -12> 2024-06-19T11:13:00.790+0000 7ff36a4d2700 1 mds.0.8058 cluster recovered. -11> 2024-06-19T11:13:00.790+0000 7ff36a4d2700 4 mds.0.8058 set_osd_epoch_barrier: epoch=33596 -10> 2024-06-19T11:13:00.790+0000 7ff3634c4700 5 mds.0.log _submit_thread 879843344432~2609 : EUpdate check_inode_max_size [metablob 0x100, 2 dirs] -9> 2024-06-19T11:13:00.791+0000 7ff3644c6700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/mds/MDCache.cc: In function 'void MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)' thread 7ff3644c6700 time 2024-06-19T11:13:00.791580+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/mds/MDCache.cc: 1660: FAILED ceph_assert(follows >= realm->get_newest_seq()) ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7ff374ad3e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7ff374ad3fdb] 3: (MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)+0x13c7) [0x55da0a7aa227] 4: (MDCache::journal_dirty_inode(MutationImpl*, EMetaBlob*, CInode*, snapid_t)+0xc5) [0x55da0a7aa3a5] 5: (Locker::check_inode_max_size(CInode*, bool, unsigned long, unsigned long, utime_t)+0x84d) [0x55da0a88ce3d] 6: (RecoveryQueue::_recovered(CInode*, int, unsigned long, utime_t)+0x4f0) [0x55da0a85ad50] 7: (MDSContext::complete(int)+0x5f) [0x55da0a9ddeef] 8: (MDSIOContextBase::complete(int)+0x524) [0x55da0a9de674] 9: (Filer::C_Probe::finish(int)+0xbb) [0x55da0aa9dc9b] 10: (Context::complete(int)+0xd) [0x55da0a6775fd] 11: (Finisher::finisher_thread_entry()+0x18d) [0x7ff374b77abd] 12: /lib64/libpthread.so.0(+0x81ca) [0x7ff3738791ca] 13: clone() -8> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client handle_log_ack log(last 7) v1 -7> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.647346+0000 mds.default.cephmon-03.chjusj (mds.0) 1 : cluster [ERR] loaded dup inode 0x10003e45d99 [415,head] v61632 at /home/balaz/.bash_history-54696.tmp, but inode 0x10003e45d99.head v61639 already exists at /home/balaz/.bash_history -6> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.648139+0000 mds.default.cephmon-03.chjusj (mds.0) 2 : cluster [ERR] loaded dup inode 0x10003e45d7c [415,head] v253612 at /home/rieder/.bash_history-10215.tmp, but inode 0x10003e45d7c.head v253630 already exists at /home/rieder/.bash_history -5> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.649483+0000 mds.default.cephmon-03.chjusj (mds.0) 3 : cluster [ERR] loaded dup inode 0x10003e45d83 [415,head] v164103 at /home/gottschling/.bash_history-44802.tmp, but inode 0x10003e45d83.head v164112 already exists at /home/gottschling/.bash_history -4> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.656221+0000 mds.default.cephmon-03.chjusj (mds.0) 4 : cluster [ERR] bad backtrace on directory inode 0x10003e42340 -3> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.737282+0000 mds.default.cephmon-03.chjusj (mds.0) 5 : cluster [ERR] bad backtrace on directory inode 0x10003e45d8b -2> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.804984+0000 mds.default.cephmon-03.chjusj (mds.0) 6 : cluster [ERR] bad backtrace on directory inode 0x10003e45d9f -1> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.805078+0000 mds.default.cephmon-03.chjusj (mds.0) 7 : cluster [ERR] bad backtrace on directory inode 0x10003e45d90 0> 2024-06-19T11:13:00.792+0000 7ff3644c6700 -1 *** Caught signal (Aborted) ** in thread 7ff3644c6700 thread_name:MR_Finisher ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7ff373883d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7ff374ad3e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7ff374ad3fdb] 6: (MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)+0x13c7) [0x55da0a7aa227] 7: (MDCache::journal_dirty_inode(MutationImpl*, EMetaBlob*, CInode*, snapid_t)+0xc5) [0x55da0a7aa3a5] 8: (Locker::check_inode_max_size(CInode*, bool, unsigned long, unsigned long, utime_t)+0x84d) [0x55da0a88ce3d] 9: (RecoveryQueue::_recovered(CInode*, int, unsigned long, utime_t)+0x4f0) [0x55da0a85ad50] 10: (MDSContext::complete(int)+0x5f) [0x55da0a9ddeef] 11: (MDSIOContextBase::complete(int)+0x524) [0x55da0a9de674] 12: (Filer::C_Probe::finish(int)+0xbb) [0x55da0aa9dc9b] 13: (Context::complete(int)+0xd) [0x55da0a6775fd] 14: (Finisher::finisher_thread_entry()+0x18d) [0x7ff374b77abd] 15: /lib64/libpthread.so.0(+0x81ca) [0x7ff3738791ca] 16: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this. --- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7ff362cc3700 / 7ff3634c4700 / md_submit 7ff363cc5700 / 7ff3644c6700 / MR_Finisher 7ff3654c8700 / PQ_Finisher 7ff365cc9700 / mds_rank_progr 7ff3664ca700 / ms_dispatch 7ff3684ce700 / ceph-mds 7ff3694d0700 / safe_timer 7ff36a4d2700 / ms_dispatch 7ff36b4d4700 / io_context_pool 7ff36c4d6700 / admin_socket 7ff36ccd7700 / msgr-worker-2 7ff36d4d8700 / msgr-worker-1 7ff36dcd9700 / msgr-worker-0 7ff375c9bb00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-03.chjusj.log --- end dump of recent events --- Any idea? Thanks Dietmar
[...]
Hi List, we are still struggeling to get our cephfs back online again, this is an update to inform you what we did so far, and we kindly ask for any input on this to get an idea on how to proceed: After resetting the journals Xiubo suggested (in a PM) to go on with the disaster recovery procedure: cephfs-data-scan init skipped creating the inodes 0x0x1 and 0x0x100 [root@ceph01-b ~]# cephfs-data-scan init Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object. We did not use --force-init and proceeded with scan_extents using a single worker, which was indeed very slow. After ~24h we interupted the scan_extents and restarted it with 32 workers which went through in about 2h15min w/o any issue. Then I started scan_inodes with 32 workers this was also finished after ~50min no output on stderr or stdout. I went on with scan_links, which after ~45 minutes threw the following error: # cephfs-data-scan scan_links Error ((2) No such file or directory) then "cephfs-data-scan cleanup" went through w/o any message and took about 9hrs 20min. Unfortunately, when starting the MDS the cephfs seems still to be in damage. I get quite some "loaded already corrupt dentry:" messages and 2 "[ERR] : bad backtrace on directory inode" errors: (In the following log I removed almost all "loaded already corrupt dentry" entries, for clarity reasons) 2024-06-23T08:06:20.934+0000 7ff05728fb00 0 set uid:gid to 167:167 (ceph:ceph) 2024-06-23T08:06:20.934+0000 7ff05728fb00 0 ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable), process ceph-mds, pid 2 2024-06-23T08:06:20.934+0000 7ff05728fb00 1 main not setting numa affinity 2024-06-23T08:06:20.934+0000 7ff05728fb00 0 pidfile_write: ignore empty --pid-file 2024-06-23T08:06:20.936+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8062 from mon.0 2024-06-23T08:06:21.583+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8063 from mon.0 2024-06-23T08:06:21.583+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Monitors have assigned me to become a standby. 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8064 from mon.0 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:standby --> up:replay 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 replay_start 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 waiting for osdmap 34327 (which blocklists prior instance) 2024-06-23T08:06:21.627+0000 7ff0452b9700 0 mds.0.cache creating system inode with ino:0x100 2024-06-23T08:06:21.627+0000 7ff0452b9700 0 mds.0.cache creating system inode with ino:0x1 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.journal EResetJournal 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe start 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe result 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe done 2024-06-23T08:06:21.656+0000 7ff045aba700 1 mds.0.8064 Finished replaying journal 2024-06-23T08:06:21.656+0000 7ff045aba700 1 mds.0.8064 making mds journal writeable 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8065 from mon.0 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:replay --> up:reconnect 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 reconnect_start 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 reopen_log 2024-06-23T08:06:22.605+0000 7ff04bac6700 1 mds.0.8064 reconnect_done 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8066 from mon.0 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:reconnect --> up:rejoin 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 rejoin_start 2024-06-23T08:06:23.609+0000 7ff04bac6700 1 mds.0.8064 rejoin_joint_start 2024-06-23T08:06:23.611+0000 7ff045aba700 1 mds.0.cache.den(0x10000000000 groups) loaded already corrupt dentry: [dentry #0x1/data/groups [bf,head] rep@0.0 NULL (dversion lock) pv=0 v= 7910497 ino=(nil) state=0 0x55cf13e9f400] 2024-06-23T08:06:23.611+0000 7ff045aba700 1 mds.0.cache.den(0x1000192ec16.1* scad_prj) loaded already corrupt dentry: [dentry #0x1/home/scad_prj [159,head] rep@0.0 NULL (dversion lock) pv=0 v=2462060 ino=(nil) state=0 0x55cf14220a00] [...] 2024-06-23T08:06:23.628+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 2024-06-23T08:06:23.668+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b [...] 2024-06-23T08:06:23.773+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawn! --- begin dump of recent events --- -9999> 2024-06-23T08:06:23.615+0000 7ff045aba700 1 mds.0.cache.den(0x1000321976d jupyterhub_slurmspawner_67002.log) loaded already corrupt dentry: [dentry #0x1/home/michelotto/jupyterhub_slurmspawner_67002.log [239,36c] rep@0.0 NULL (dversion lock) pv=0 v=26855917 ino=(nil) state=0 0x55cf147ff680] [...] -9878> 2024-06-23T08:06:23.616+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf13f13c00 auth_method 0 [...] -9877> 2024-06-23T08:06:23.616+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf146d1800 auth_method 0 -9458> 2024-06-23T08:06:23.620+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f13400 auth_method 0 -9353> 2024-06-23T08:06:23.622+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf146d0800 auth_method 0 -8980> 2024-06-23T08:06:23.625+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14d9f400 auth_method 0 -8978> 2024-06-23T08:06:23.625+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf14d9fc00 auth_method 0 -8849> 2024-06-23T08:06:23.628+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 -8574> 2024-06-23T08:06:23.633+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14323400 auth_method 0 -8570> 2024-06-23T08:06:23.633+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf1485e400 auth_method 0 -8564> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14322c00 auth_method 0 -8561> 2024-06-23T08:06:23.633+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f12800 auth_method 0 -8555> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14f48000 auth_method 0 -8546> 2024-06-23T08:06:23.633+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf1485f800 auth_method 0 -8541> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf1516b000 auth_method 0 -8470> 2024-06-23T08:06:23.634+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14322400 auth_method 0 -8451> 2024-06-23T08:06:23.635+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f13800 auth_method 0 -8445> 2024-06-23T08:06:23.635+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf1485e800 auth_method 0 -8243> 2024-06-23T08:06:23.637+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf141ce400 auth_method 0 -7381> 2024-06-23T08:06:23.645+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf1485f400 auth_method 0 -6319> 2024-06-23T08:06:23.660+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14f48400 auth_method 0 -5946> 2024-06-23T08:06:23.666+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf146d1400 auth_method 0 [...] -5677> 2024-06-23T08:06:23.668+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b [...] -8> 2024-06-23T08:06:23.753+0000 7ff045aba700 5 mds.beacon.default.cephmon-01.cepqjp set_want_state: up:rejoin -> down:damaged -7> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client log_queue is 2 last_log 2 sent 0 num 2 unsent 2 sending 2 -6> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client will send 2024-06-23T08:06:23.629743+0000 mds.default.cephmon-01.cepqjp (mds.0) 1 : cluster [ERR] bad backtrace on directory inode 0x10003e42340 -5> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client will send 2024-06-23T08:06:23.669673+0000 mds.default.cephmon-01.cepqjp (mds.0) 2 : cluster [ERR] bad backtrace on directory inode 0x10003e45d8b -4> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 monclient: _send_mon_message to mon.cephmon-01 at v2:10.1.3.21:3300/0 -3> 2024-06-23T08:06:23.753+0000 7ff045aba700 5 mds.beacon.default.cephmon-01.cepqjp Sending beacon down:damaged seq 4 -2> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 monclient: _send_mon_message to mon.cephmon-01 at v2:10.1.3.21:3300/0 -1> 2024-06-23T08:06:23.773+0000 7ff04e2cb700 5 mds.beacon.default.cephmon-01.cepqjp received beacon reply down:damaged seq 4 rtt 0.0200001 0> 2024-06-23T08:06:23.773+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawn! --- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7ff045aba700 / MR_Finisher 7ff04e2cb700 / msgr-worker-2 7ff04eacc700 / msgr-worker-1 7ff04f2cd700 / msgr-worker-0 max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-01.cepqjp.log --- end dump of recent events --- 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp e: '/usr/bin/ceph-mds' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 0: '/usr/bin/ceph-mds' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 1: '-n' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 2: 'mds.default.cephmon-01.cepqjp' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 3: '-f' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 4: '--setuser' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 5: 'ceph' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 6: '--setgroup' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 7: 'ceph' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 8: '--default-log-to-file=false' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 9: '--default-log-to-journald=true' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 10: '--default-log-to-stderr=false' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawning with exe /usr/bin/ceph-mds 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp exe_path /proc/self/exe 2024-06-23T08:06:23.812+0000 7fe58d619b00 0 ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable), process ceph-mds, pid 2 2024-06-23T08:06:23.812+0000 7fe58d619b00 1 main not setting numa affinity 2024-06-23T08:06:23.813+0000 7fe58d619b00 0 pidfile_write: ignore empty --pid-file 2024-06-23T08:06:23.814+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8067 from mon.0 2024-06-23T08:06:24.772+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8068 from mon.0 2024-06-23T08:06:24.772+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Monitors have assigned me to become a standby. 2024-06-23T08:49:28.778+0000 7fe584272700 1 mds.default.cephmon-01.cepqjp asok_command: heap {heapcmd=stats,prefix=heap} (starting...) 2024-06-23T22:00:04.664+0000 7fe583a71700 -1 received signal: Hangup from Kernel ( Could be generated by pthread_kill(), raise(), abort(), alarm() ) UID: 0 Any ideas how to proceed? Would rerunning the cephfs-data-scan sequence do any harm or give us a chance to resolve this? Would removing "snapBackup_head" omap keys help to fix the bad backtrace error? [ERR] : bad backtrace on directory inode 0x10003e42340 This are the corresponding omapvals: # rados --cluster ceph -p ssd-rep-metadata-pool listomapvals 10003e42340.00000000 snapBackup_head value (484 bytes) : 00000000 23 04 00 00 00 00 00 00 49 13 06 b9 01 00 00 41 |#.......I......A| 00000010 23 e4 03 00 01 00 00 00 00 00 00 a0 70 72 66 c6 |#...........prf.| 00000020 ba 9f 32 ed 41 00 00 07 9d 00 00 50 c3 00 00 01 |..2.A......P....| 00000030 00 00 00 00 02 00 00 00 00 00 00 00 02 02 18 00 |................| 00000040 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ff ff |................| 00000050 ff ff ff ff ff ff 00 00 00 00 00 00 00 00 00 00 |................| 00000060 00 00 01 00 00 00 ff ff ff ff ff ff ff ff 00 00 |................| 00000070 00 00 00 00 00 00 00 00 00 00 a0 70 72 66 c6 ba |...........prf..| 00000080 9f 32 a0 70 72 66 75 78 90 32 00 00 00 00 00 00 |.2.prfux.2......| 00000090 00 00 03 02 28 00 00 00 00 00 00 00 00 00 00 00 |....(...........| 000000a0 a0 70 72 66 c6 ba 9f 32 01 00 00 00 00 00 00 00 |.prf...2........| 000000b0 00 00 00 00 00 00 00 00 01 00 00 00 00 00 00 00 |................| 000000c0 03 02 38 00 00 00 00 00 00 00 00 00 00 00 b6 16 |..8.............| 000000d0 00 00 00 00 00 00 01 00 00 00 00 00 00 00 01 00 |................| 000000e0 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 000000f0 00 00 00 00 00 00 2a 74 72 66 bc b9 6c 07 03 02 |......*trf..l...| 00000100 38 00 00 00 00 00 00 00 00 00 00 00 b6 16 00 00 |8...............| 00000110 00 00 00 00 01 00 00 00 00 00 00 00 01 00 00 00 |................| 00000120 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 00000130 00 00 00 00 2a 74 72 66 bc b9 6c 07 26 05 00 00 |....*trf..l.&...| 00000140 00 00 00 00 00 00 00 00 00 00 00 00 01 00 00 00 |................| 00000150 00 00 00 00 02 00 00 00 00 00 00 00 00 00 00 00 |................| 00000160 00 00 00 00 00 00 00 00 ff ff ff ff ff ff ff ff |................| 00000170 00 00 00 00 01 01 10 00 00 00 00 00 00 00 00 00 |................| 00000180 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 00000190 00 00 00 00 00 00 00 00 00 00 00 00 00 00 a0 70 |...............p| 000001a0 72 66 75 78 90 32 01 00 00 00 00 00 00 00 ff ff |rfux.2..........| 000001b0 ff ff 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 000001c0 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 000001d0 00 00 00 00 00 00 00 00 fe ff ff ff ff ff ff ff |................| 000001e0 00 00 00 00 |....| 000001e4 Thanks for any help Dietmar On 6/19/24 13:42, Dietmar Rieder wrote:
On 6/19/24 11:15, Dietmar Rieder wrote:
On 6/19/24 10:30, Xiubo Li wrote:
On 6/19/24 16:13, Dietmar Rieder wrote:
Hi Xiubo,
[...]
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
This is a known bug, please see https://tracker.ceph.com/issues/61009.
As a workaround I am afraid you need to trim the journal logs first and then try to restart the MDS daemons, And at the same time please follow the workaround in https://tracker.ceph.com/issues/61009#note-26
I see, I'll try to do this. Are there any caveats or issues to expect by trimming the journal logs?
Certainly you will lose the dirty metadata in the journals.
Is there a step by step guide on how to perform the trimming? Should all MDS be stopped before?
Please follow https://docs.ceph.com/en/nautilus/cephfs/disaster-recovery-experts/#disaster....
OK, when I run the cephfs-journal-tool I get an error:
# cephfs-journal-tool journal export backup.bin Error ((22) Invalid argument)
My cluster is managed by caphadm, so (in my stress situation) I'm not able find the correct way to use cephfs-journal-tool
I'm sure it is something stupid that I'm missing but I'd be happy for any hint.
I ran the disaster recovery procedures now, as follows:
[root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 event recover_dentries summary Events by type: OPEN: 8737 PURGED: 1 SESSION: 9 SESSIONS: 2 SUBTREEMAP: 128 TABLECLIENT: 2 TABLESERVER: 30 UPDATE: 9207 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 event recover_dentries summary Events by type: OPEN: 3 SESSION: 1 SUBTREEMAP: 34 UPDATE: 32965 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 event recover_dentries summary Events by type: OPEN: 5289 SESSION: 10 SESSIONS: 3 SUBTREEMAP: 128 UPDATE: 76448 Errors: 0
[root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:all journal inspect Overall journal integrity: OK Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 journal inspect Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 journal inspect Overall journal integrity: OK
[root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 journal reset old journal was 879331755046~508520587 new journal start will be 879843344384 (3068751 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal reset old journal was 934711229813~120432327 new journal start will be 934834864128 (3201988 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 journal reset old journal was 1334153584288~252692691 new journal start will be 1334409428992 (3152013 bytes past old end) writing journal head writing EResetJournal entry done
[root@ceph01-b /]# cephfs-table-tool all reset session { "0": { "data": {}, "result": 0 }, "1": { "data": {}, "result": 0 }, "2": { "data": {}, "result": 0 } }
[root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: OK
[root@ceph01-b /]# ceph fs reset cephfs --yes-i-really-mean-it
But now I hit the error below:
-20> 2024-06-19T11:13:00.610+0000 7ff3694d0700 10 monclient: _send_mon_message to mon.cephmon-03 at v2:10.1.3.23:3300/0 -19> 2024-06-19T11:13:00.637+0000 7ff3664ca700 2 mds.0.cache Memory usage: total 485928, rss 170860, heap 207156, baseline 182580, 0 / 33434 inodes have caps, 0 caps, 0 caps per inode -18> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.default.cephmon-03.chjusj Updating MDS map to version 8061 from mon.1 -17> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 handle_mds_map i am now mds.0.8058 -16> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 handle_mds_map state change up:rejoin --> up:active -15> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 recovery_done -- successful recovery! -14> 2024-06-19T11:13:00.788+0000 7ff36a4d2700 1 mds.0.8058 active_start -13> 2024-06-19T11:13:00.789+0000 7ff36dcd9700 5 mds.beacon.default.cephmon-03.chjusj received beacon reply up:active seq 4 rtt 0.955007 -12> 2024-06-19T11:13:00.790+0000 7ff36a4d2700 1 mds.0.8058 cluster recovered. -11> 2024-06-19T11:13:00.790+0000 7ff36a4d2700 4 mds.0.8058 set_osd_epoch_barrier: epoch=33596 -10> 2024-06-19T11:13:00.790+0000 7ff3634c4700 5 mds.0.log _submit_thread 879843344432~2609 : EUpdate check_inode_max_size [metablob 0x100, 2 dirs] -9> 2024-06-19T11:13:00.791+0000 7ff3644c6700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/mds/MDCache.cc: In function 'void MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)' thread 7ff3644c6700 time 2024-06-19T11:13:00.791580+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/mds/MDCache.cc: 1660: FAILED ceph_assert(follows >= realm->get_newest_seq())
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7ff374ad3e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7ff374ad3fdb] 3: (MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)+0x13c7) [0x55da0a7aa227] 4: (MDCache::journal_dirty_inode(MutationImpl*, EMetaBlob*, CInode*, snapid_t)+0xc5) [0x55da0a7aa3a5] 5: (Locker::check_inode_max_size(CInode*, bool, unsigned long, unsigned long, utime_t)+0x84d) [0x55da0a88ce3d] 6: (RecoveryQueue::_recovered(CInode*, int, unsigned long, utime_t)+0x4f0) [0x55da0a85ad50] 7: (MDSContext::complete(int)+0x5f) [0x55da0a9ddeef] 8: (MDSIOContextBase::complete(int)+0x524) [0x55da0a9de674] 9: (Filer::C_Probe::finish(int)+0xbb) [0x55da0aa9dc9b] 10: (Context::complete(int)+0xd) [0x55da0a6775fd] 11: (Finisher::finisher_thread_entry()+0x18d) [0x7ff374b77abd] 12: /lib64/libpthread.so.0(+0x81ca) [0x7ff3738791ca] 13: clone()
-8> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client handle_log_ack log(last 7) v1 -7> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.647346+0000 mds.default.cephmon-03.chjusj (mds.0) 1 : cluster [ERR] loaded dup inode 0x10003e45d99 [415,head] v61632 at /home/balaz/.bash_history-54696.tmp, but inode 0x10003e45d99.head v61639 already exists at /home/balaz/.bash_history -6> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.648139+0000 mds.default.cephmon-03.chjusj (mds.0) 2 : cluster [ERR] loaded dup inode 0x10003e45d7c [415,head] v253612 at /home/rieder/.bash_history-10215.tmp, but inode 0x10003e45d7c.head v253630 already exists at /home/rieder/.bash_history -5> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.649483+0000 mds.default.cephmon-03.chjusj (mds.0) 3 : cluster [ERR] loaded dup inode 0x10003e45d83 [415,head] v164103 at /home/gottschling/.bash_history-44802.tmp, but inode 0x10003e45d83.head v164112 already exists at /home/gottschling/.bash_history -4> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.656221+0000 mds.default.cephmon-03.chjusj (mds.0) 4 : cluster [ERR] bad backtrace on directory inode 0x10003e42340 -3> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.737282+0000 mds.default.cephmon-03.chjusj (mds.0) 5 : cluster [ERR] bad backtrace on directory inode 0x10003e45d8b -2> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.804984+0000 mds.default.cephmon-03.chjusj (mds.0) 6 : cluster [ERR] bad backtrace on directory inode 0x10003e45d9f -1> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.805078+0000 mds.default.cephmon-03.chjusj (mds.0) 7 : cluster [ERR] bad backtrace on directory inode 0x10003e45d90 0> 2024-06-19T11:13:00.792+0000 7ff3644c6700 -1 *** Caught signal (Aborted) ** in thread 7ff3644c6700 thread_name:MR_Finisher
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7ff373883d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7ff374ad3e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7ff374ad3fdb] 6: (MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)+0x13c7) [0x55da0a7aa227] 7: (MDCache::journal_dirty_inode(MutationImpl*, EMetaBlob*, CInode*, snapid_t)+0xc5) [0x55da0a7aa3a5] 8: (Locker::check_inode_max_size(CInode*, bool, unsigned long, unsigned long, utime_t)+0x84d) [0x55da0a88ce3d] 9: (RecoveryQueue::_recovered(CInode*, int, unsigned long, utime_t)+0x4f0) [0x55da0a85ad50] 10: (MDSContext::complete(int)+0x5f) [0x55da0a9ddeef] 11: (MDSIOContextBase::complete(int)+0x524) [0x55da0a9de674] 12: (Filer::C_Probe::finish(int)+0xbb) [0x55da0aa9dc9b] 13: (Context::complete(int)+0xd) [0x55da0a6775fd] 14: (Finisher::finisher_thread_entry()+0x18d) [0x7ff374b77abd] 15: /lib64/libpthread.so.0(+0x81ca) [0x7ff3738791ca] 16: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7ff362cc3700 / 7ff3634c4700 / md_submit 7ff363cc5700 / 7ff3644c6700 / MR_Finisher 7ff3654c8700 / PQ_Finisher 7ff365cc9700 / mds_rank_progr 7ff3664ca700 / ms_dispatch 7ff3684ce700 / ceph-mds 7ff3694d0700 / safe_timer 7ff36a4d2700 / ms_dispatch 7ff36b4d4700 / io_context_pool 7ff36c4d6700 / admin_socket 7ff36ccd7700 / msgr-worker-2 7ff36d4d8700 / msgr-worker-1 7ff36dcd9700 / msgr-worker-0 7ff375c9bb00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-03.chjusj.log --- end dump of recent events ---
Any idea?
Thanks
Dietmar
[...]
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
(resending this, the original message seems that it didn't make it through between all the SPAM recently sent to the list, my apologies if it doubles at some point) Hi List, we are still struggeling to get our cephfs back online again, this is an update to inform you what we did so far, and we kindly ask for any input on this to get an idea on how to proceed: After resetting the journals Xiubo suggested (in a PM) to go on with the disaster recovery procedure: cephfs-data-scan init skipped creating the inodes 0x0x1 and 0x0x100 [root@ceph01-b ~]# cephfs-data-scan init Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object. We did not use --force-init and proceeded with scan_extents using a single worker, which was indeed very slow. After ~24h we interupted the scan_extents and restarted it with 32 workers which went through in about 2h15min w/o any issue. Then I started scan_inodes with 32 workers this was also finished after ~50min no output on stderr or stdout. I went on with scan_links, which after ~45 minutes threw the following error: # cephfs-data-scan scan_links Error ((2) No such file or directory) then "cephfs-data-scan cleanup" went through w/o any message and took about 9hrs 20min. Unfortunately, when starting the MDS the cephfs seems still to be in damage. I get quite some "loaded already corrupt dentry:" messages and 2 "[ERR] : bad backtrace on directory inode" errors: (In the following log I removed almost all "loaded already corrupt dentry" entries, for clarity reasons) 2024-06-23T08:06:20.934+0000 7ff05728fb00 0 set uid:gid to 167:167 (ceph:ceph) 2024-06-23T08:06:20.934+0000 7ff05728fb00 0 ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable), process ceph-mds, pid 2 2024-06-23T08:06:20.934+0000 7ff05728fb00 1 main not setting numa affinity 2024-06-23T08:06:20.934+0000 7ff05728fb00 0 pidfile_write: ignore empty --pid-file 2024-06-23T08:06:20.936+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8062 from mon.0 2024-06-23T08:06:21.583+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8063 from mon.0 2024-06-23T08:06:21.583+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Monitors have assigned me to become a standby. 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8064 from mon.0 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:standby --> up:replay 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 replay_start 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 waiting for osdmap 34327 (which blocklists prior instance) 2024-06-23T08:06:21.627+0000 7ff0452b9700 0 mds.0.cache creating system inode with ino:0x100 2024-06-23T08:06:21.627+0000 7ff0452b9700 0 mds.0.cache creating system inode with ino:0x1 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.journal EResetJournal 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe start 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe result 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe done 2024-06-23T08:06:21.656+0000 7ff045aba700 1 mds.0.8064 Finished replaying journal 2024-06-23T08:06:21.656+0000 7ff045aba700 1 mds.0.8064 making mds journal writeable 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8065 from mon.0 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:replay --> up:reconnect 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 reconnect_start 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 reopen_log 2024-06-23T08:06:22.605+0000 7ff04bac6700 1 mds.0.8064 reconnect_done 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8066 from mon.0 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:reconnect --> up:rejoin 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 rejoin_start 2024-06-23T08:06:23.609+0000 7ff04bac6700 1 mds.0.8064 rejoin_joint_start 2024-06-23T08:06:23.611+0000 7ff045aba700 1 mds.0.cache.den(0x10000000000 groups) loaded already corrupt dentry: [dentry #0x1/data/groups [bf,head] rep@0.0 NULL (dversion lock) pv=0 v= 7910497 ino=(nil) state=0 0x55cf13e9f400] 2024-06-23T08:06:23.611+0000 7ff045aba700 1 mds.0.cache.den(0x1000192ec16.1* scad_prj) loaded already corrupt dentry: [dentry #0x1/home/scad_prj [159,head] rep@0.0 NULL (dversion lock) pv=0 v=2462060 ino=(nil) state=0 0x55cf14220a00] [...] 2024-06-23T08:06:23.628+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 2024-06-23T08:06:23.668+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b [...] 2024-06-23T08:06:23.773+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawn! --- begin dump of recent events --- -9999> 2024-06-23T08:06:23.615+0000 7ff045aba700 1 mds.0.cache.den(0x1000321976d jupyterhub_slurmspawner_67002.log) loaded already corrupt dentry: [dentry #0x1/home/michelotto/jupyterhub_slurmspawner_67002.log [239,36c] rep@0.0 NULL (dversion lock) pv=0 v=26855917 ino=(nil) state=0 0x55cf147ff680] [...] -9878> 2024-06-23T08:06:23.616+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf13f13c00 auth_method 0 [...] -9877> 2024-06-23T08:06:23.616+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf146d1800 auth_method 0 -9458> 2024-06-23T08:06:23.620+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f13400 auth_method 0 -9353> 2024-06-23T08:06:23.622+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf146d0800 auth_method 0 -8980> 2024-06-23T08:06:23.625+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14d9f400 auth_method 0 -8978> 2024-06-23T08:06:23.625+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf14d9fc00 auth_method 0 -8849> 2024-06-23T08:06:23.628+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 -8574> 2024-06-23T08:06:23.633+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14323400 auth_method 0 -8570> 2024-06-23T08:06:23.633+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf1485e400 auth_method 0 -8564> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14322c00 auth_method 0 -8561> 2024-06-23T08:06:23.633+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f12800 auth_method 0 -8555> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14f48000 auth_method 0 -8546> 2024-06-23T08:06:23.633+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf1485f800 auth_method 0 -8541> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf1516b000 auth_method 0 -8470> 2024-06-23T08:06:23.634+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14322400 auth_method 0 -8451> 2024-06-23T08:06:23.635+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f13800 auth_method 0 -8445> 2024-06-23T08:06:23.635+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf1485e800 auth_method 0 -8243> 2024-06-23T08:06:23.637+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf141ce400 auth_method 0 -7381> 2024-06-23T08:06:23.645+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf1485f400 auth_method 0 -6319> 2024-06-23T08:06:23.660+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14f48400 auth_method 0 -5946> 2024-06-23T08:06:23.666+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf146d1400 auth_method 0 [...] -5677> 2024-06-23T08:06:23.668+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b [...] -8> 2024-06-23T08:06:23.753+0000 7ff045aba700 5 mds.beacon.default.cephmon-01.cepqjp set_want_state: up:rejoin -> down:damaged -7> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client log_queue is 2 last_log 2 sent 0 num 2 unsent 2 sending 2 -6> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client will send 2024-06-23T08:06:23.629743+0000 mds.default.cephmon-01.cepqjp (mds.0) 1 : cluster [ERR] bad backtrace on directory inode 0x10003e42340 -5> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client will send 2024-06-23T08:06:23.669673+0000 mds.default.cephmon-01.cepqjp (mds.0) 2 : cluster [ERR] bad backtrace on directory inode 0x10003e45d8b -4> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 monclient: _send_mon_message to mon.cephmon-01 at v2:10.1.3.21:3300/0 -3> 2024-06-23T08:06:23.753+0000 7ff045aba700 5 mds.beacon.default.cephmon-01.cepqjp Sending beacon down:damaged seq 4 -2> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 monclient: _send_mon_message to mon.cephmon-01 at v2:10.1.3.21:3300/0 -1> 2024-06-23T08:06:23.773+0000 7ff04e2cb700 5 mds.beacon.default.cephmon-01.cepqjp received beacon reply down:damaged seq 4 rtt 0.0200001 0> 2024-06-23T08:06:23.773+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawn! --- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7ff045aba700 / MR_Finisher 7ff04e2cb700 / msgr-worker-2 7ff04eacc700 / msgr-worker-1 7ff04f2cd700 / msgr-worker-0 max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-01.cepqjp.log --- end dump of recent events --- 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp e: '/usr/bin/ceph-mds' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 0: '/usr/bin/ceph-mds' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 1: '-n' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 2: 'mds.default.cephmon-01.cepqjp' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 3: '-f' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 4: '--setuser' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 5: 'ceph' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 6: '--setgroup' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 7: 'ceph' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 8: '--default-log-to-file=false' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 9: '--default-log-to-journald=true' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 10: '--default-log-to-stderr=false' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawning with exe /usr/bin/ceph-mds 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp exe_path /proc/self/exe 2024-06-23T08:06:23.812+0000 7fe58d619b00 0 ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable), process ceph-mds, pid 2 2024-06-23T08:06:23.812+0000 7fe58d619b00 1 main not setting numa affinity 2024-06-23T08:06:23.813+0000 7fe58d619b00 0 pidfile_write: ignore empty --pid-file 2024-06-23T08:06:23.814+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8067 from mon.0 2024-06-23T08:06:24.772+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8068 from mon.0 2024-06-23T08:06:24.772+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Monitors have assigned me to become a standby. 2024-06-23T08:49:28.778+0000 7fe584272700 1 mds.default.cephmon-01.cepqjp asok_command: heap {heapcmd=stats,prefix=heap} (starting...) 2024-06-23T22:00:04.664+0000 7fe583a71700 -1 received signal: Hangup from Kernel ( Could be generated by pthread_kill(), raise(), abort(), alarm() ) UID: 0 Any ideas how to proceed? Would rerunning the cephfs-data-scan sequence do any harm or give us a chance to resolve this? Would removing "snapBackup_head" omap keys help to fix the bad backtrace error? [ERR] : bad backtrace on directory inode 0x10003e42340 This are the corresponding omapvals: # rados --cluster ceph -p ssd-rep-metadata-pool listomapvals 10003e42340.00000000 snapBackup_head value (484 bytes) : 00000000 23 04 00 00 00 00 00 00 49 13 06 b9 01 00 00 41 |#.......I......A| 00000010 23 e4 03 00 01 00 00 00 00 00 00 a0 70 72 66 c6 |#...........prf.| 00000020 ba 9f 32 ed 41 00 00 07 9d 00 00 50 c3 00 00 01 |..2.A......P....| 00000030 00 00 00 00 02 00 00 00 00 00 00 00 02 02 18 00 |................| 00000040 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ff ff |................| 00000050 ff ff ff ff ff ff 00 00 00 00 00 00 00 00 00 00 |................| 00000060 00 00 01 00 00 00 ff ff ff ff ff ff ff ff 00 00 |................| 00000070 00 00 00 00 00 00 00 00 00 00 a0 70 72 66 c6 ba |...........prf..| 00000080 9f 32 a0 70 72 66 75 78 90 32 00 00 00 00 00 00 |.2.prfux.2......| 00000090 00 00 03 02 28 00 00 00 00 00 00 00 00 00 00 00 |....(...........| 000000a0 a0 70 72 66 c6 ba 9f 32 01 00 00 00 00 00 00 00 |.prf...2........| 000000b0 00 00 00 00 00 00 00 00 01 00 00 00 00 00 00 00 |................| 000000c0 03 02 38 00 00 00 00 00 00 00 00 00 00 00 b6 16 |..8.............| 000000d0 00 00 00 00 00 00 01 00 00 00 00 00 00 00 01 00 |................| 000000e0 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 000000f0 00 00 00 00 00 00 2a 74 72 66 bc b9 6c 07 03 02 |......*trf..l...| 00000100 38 00 00 00 00 00 00 00 00 00 00 00 b6 16 00 00 |8...............| 00000110 00 00 00 00 01 00 00 00 00 00 00 00 01 00 00 00 |................| 00000120 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 00000130 00 00 00 00 2a 74 72 66 bc b9 6c 07 26 05 00 00 |....*trf..l.&...| 00000140 00 00 00 00 00 00 00 00 00 00 00 00 01 00 00 00 |................| 00000150 00 00 00 00 02 00 00 00 00 00 00 00 00 00 00 00 |................| 00000160 00 00 00 00 00 00 00 00 ff ff ff ff ff ff ff ff |................| 00000170 00 00 00 00 01 01 10 00 00 00 00 00 00 00 00 00 |................| 00000180 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 00000190 00 00 00 00 00 00 00 00 00 00 00 00 00 00 a0 70 |...............p| 000001a0 72 66 75 78 90 32 01 00 00 00 00 00 00 00 ff ff |rfux.2..........| 000001b0 ff ff 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 000001c0 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 |................| 000001d0 00 00 00 00 00 00 00 00 fe ff ff ff ff ff ff ff |................| 000001e0 00 00 00 00 |....| 000001e4 Thanks for any help Dietmar On 6/19/24 13:42, Dietmar Rieder wrote:
On 6/19/24 11:15, Dietmar Rieder wrote:
On 6/19/24 10:30, Xiubo Li wrote:
On 6/19/24 16:13, Dietmar Rieder wrote:
Hi Xiubo,
[...]
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
This is a known bug, please see https://tracker.ceph.com/issues/61009.
As a workaround I am afraid you need to trim the journal logs first and then try to restart the MDS daemons, And at the same time please follow the workaround in https://tracker.ceph.com/issues/61009#note-26
I see, I'll try to do this. Are there any caveats or issues to expect by trimming the journal logs?
Certainly you will lose the dirty metadata in the journals.
Is there a step by step guide on how to perform the trimming? Should all MDS be stopped before?
Please follow https://docs.ceph.com/en/nautilus/cephfs/disaster-recovery-experts/#disaster....
OK, when I run the cephfs-journal-tool I get an error:
# cephfs-journal-tool journal export backup.bin Error ((22) Invalid argument)
My cluster is managed by caphadm, so (in my stress situation) I'm not able find the correct way to use cephfs-journal-tool
I'm sure it is something stupid that I'm missing but I'd be happy for any hint.
I ran the disaster recovery procedures now, as follows:
[root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 event recover_dentries summary Events by type: OPEN: 8737 PURGED: 1 SESSION: 9 SESSIONS: 2 SUBTREEMAP: 128 TABLECLIENT: 2 TABLESERVER: 30 UPDATE: 9207 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 event recover_dentries summary Events by type: OPEN: 3 SESSION: 1 SUBTREEMAP: 34 UPDATE: 32965 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 event recover_dentries summary Events by type: OPEN: 5289 SESSION: 10 SESSIONS: 3 SUBTREEMAP: 128 UPDATE: 76448 Errors: 0
[root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:all journal inspect Overall journal integrity: OK Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 journal inspect Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 journal inspect Overall journal integrity: OK
[root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 journal reset old journal was 879331755046~508520587 new journal start will be 879843344384 (3068751 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal reset old journal was 934711229813~120432327 new journal start will be 934834864128 (3201988 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 journal reset old journal was 1334153584288~252692691 new journal start will be 1334409428992 (3152013 bytes past old end) writing journal head writing EResetJournal entry done
[root@ceph01-b /]# cephfs-table-tool all reset session { "0": { "data": {}, "result": 0 }, "1": { "data": {}, "result": 0 }, "2": { "data": {}, "result": 0 } }
[root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: OK
[root@ceph01-b /]# ceph fs reset cephfs --yes-i-really-mean-it
But now I hit the error below:
-20> 2024-06-19T11:13:00.610+0000 7ff3694d0700 10 monclient: _send_mon_message to mon.cephmon-03 at v2:10.1.3.23:3300/0 -19> 2024-06-19T11:13:00.637+0000 7ff3664ca700 2 mds.0.cache Memory usage: total 485928, rss 170860, heap 207156, baseline 182580, 0 / 33434 inodes have caps, 0 caps, 0 caps per inode -18> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.default.cephmon-03.chjusj Updating MDS map to version 8061 from mon.1 -17> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 handle_mds_map i am now mds.0.8058 -16> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 handle_mds_map state change up:rejoin --> up:active -15> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 recovery_done -- successful recovery! -14> 2024-06-19T11:13:00.788+0000 7ff36a4d2700 1 mds.0.8058 active_start -13> 2024-06-19T11:13:00.789+0000 7ff36dcd9700 5 mds.beacon.default.cephmon-03.chjusj received beacon reply up:active seq 4 rtt 0.955007 -12> 2024-06-19T11:13:00.790+0000 7ff36a4d2700 1 mds.0.8058 cluster recovered. -11> 2024-06-19T11:13:00.790+0000 7ff36a4d2700 4 mds.0.8058 set_osd_epoch_barrier: epoch=33596 -10> 2024-06-19T11:13:00.790+0000 7ff3634c4700 5 mds.0.log _submit_thread 879843344432~2609 : EUpdate check_inode_max_size [metablob 0x100, 2 dirs] -9> 2024-06-19T11:13:00.791+0000 7ff3644c6700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/mds/MDCache.cc: In function 'void MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)' thread 7ff3644c6700 time 2024-06-19T11:13:00.791580+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/mds/MDCache.cc: 1660: FAILED ceph_assert(follows >= realm->get_newest_seq())
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7ff374ad3e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7ff374ad3fdb] 3: (MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)+0x13c7) [0x55da0a7aa227] 4: (MDCache::journal_dirty_inode(MutationImpl*, EMetaBlob*, CInode*, snapid_t)+0xc5) [0x55da0a7aa3a5] 5: (Locker::check_inode_max_size(CInode*, bool, unsigned long, unsigned long, utime_t)+0x84d) [0x55da0a88ce3d] 6: (RecoveryQueue::_recovered(CInode*, int, unsigned long, utime_t)+0x4f0) [0x55da0a85ad50] 7: (MDSContext::complete(int)+0x5f) [0x55da0a9ddeef] 8: (MDSIOContextBase::complete(int)+0x524) [0x55da0a9de674] 9: (Filer::C_Probe::finish(int)+0xbb) [0x55da0aa9dc9b] 10: (Context::complete(int)+0xd) [0x55da0a6775fd] 11: (Finisher::finisher_thread_entry()+0x18d) [0x7ff374b77abd] 12: /lib64/libpthread.so.0(+0x81ca) [0x7ff3738791ca] 13: clone()
-8> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client handle_log_ack log(last 7) v1 -7> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.647346+0000 mds.default.cephmon-03.chjusj (mds.0) 1 : cluster [ERR] loaded dup inode 0x10003e45d99 [415,head] v61632 at /home/balaz/.bash_history-54696.tmp, but inode 0x10003e45d99.head v61639 already exists at /home/balaz/.bash_history -6> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.648139+0000 mds.default.cephmon-03.chjusj (mds.0) 2 : cluster [ERR] loaded dup inode 0x10003e45d7c [415,head] v253612 at /home/rieder/.bash_history-10215.tmp, but inode 0x10003e45d7c.head v253630 already exists at /home/rieder/.bash_history -5> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.649483+0000 mds.default.cephmon-03.chjusj (mds.0) 3 : cluster [ERR] loaded dup inode 0x10003e45d83 [415,head] v164103 at /home/gottschling/.bash_history-44802.tmp, but inode 0x10003e45d83.head v164112 already exists at /home/gottschling/.bash_history -4> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.656221+0000 mds.default.cephmon-03.chjusj (mds.0) 4 : cluster [ERR] bad backtrace on directory inode 0x10003e42340 -3> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.737282+0000 mds.default.cephmon-03.chjusj (mds.0) 5 : cluster [ERR] bad backtrace on directory inode 0x10003e45d8b -2> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.804984+0000 mds.default.cephmon-03.chjusj (mds.0) 6 : cluster [ERR] bad backtrace on directory inode 0x10003e45d9f -1> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.805078+0000 mds.default.cephmon-03.chjusj (mds.0) 7 : cluster [ERR] bad backtrace on directory inode 0x10003e45d90 0> 2024-06-19T11:13:00.792+0000 7ff3644c6700 -1 *** Caught signal (Aborted) ** in thread 7ff3644c6700 thread_name:MR_Finisher
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7ff373883d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7ff374ad3e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7ff374ad3fdb] 6: (MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)+0x13c7) [0x55da0a7aa227] 7: (MDCache::journal_dirty_inode(MutationImpl*, EMetaBlob*, CInode*, snapid_t)+0xc5) [0x55da0a7aa3a5] 8: (Locker::check_inode_max_size(CInode*, bool, unsigned long, unsigned long, utime_t)+0x84d) [0x55da0a88ce3d] 9: (RecoveryQueue::_recovered(CInode*, int, unsigned long, utime_t)+0x4f0) [0x55da0a85ad50] 10: (MDSContext::complete(int)+0x5f) [0x55da0a9ddeef] 11: (MDSIOContextBase::complete(int)+0x524) [0x55da0a9de674] 12: (Filer::C_Probe::finish(int)+0xbb) [0x55da0aa9dc9b] 13: (Context::complete(int)+0xd) [0x55da0a6775fd] 14: (Finisher::finisher_thread_entry()+0x18d) [0x7ff374b77abd] 15: /lib64/libpthread.so.0(+0x81ca) [0x7ff3738791ca] 16: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7ff362cc3700 / 7ff3634c4700 / md_submit 7ff363cc5700 / 7ff3644c6700 / MR_Finisher 7ff3654c8700 / PQ_Finisher 7ff365cc9700 / mds_rank_progr 7ff3664ca700 / ms_dispatch 7ff3684ce700 / ceph-mds 7ff3694d0700 / safe_timer 7ff36a4d2700 / ms_dispatch 7ff36b4d4700 / io_context_pool 7ff36c4d6700 / admin_socket 7ff36ccd7700 / msgr-worker-2 7ff36d4d8700 / msgr-worker-1 7ff36dcd9700 / msgr-worker-0 7ff375c9bb00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-03.chjusj.log --- end dump of recent events ---
Any idea?
Thanks
Dietmar
[...]
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at -- _______________________________________________ D i e t m a r R i e d e r, Mag.Dr. Head of Bioinformatics Core Facility Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 Mobile: +43 676 8716 72402 Fax: +43 512 9003 74400 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
On Mon, Jun 24, 2024 at 5:22 PM Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
(resending this, the original message seems that it didn't make it through between all the SPAM recently sent to the list, my apologies if it doubles at some point)
Hi List,
we are still struggeling to get our cephfs back online again, this is an update to inform you what we did so far, and we kindly ask for any input on this to get an idea on how to proceed:
After resetting the journals Xiubo suggested (in a PM) to go on with the disaster recovery procedure:
cephfs-data-scan init skipped creating the inodes 0x0x1 and 0x0x100
[root@ceph01-b ~]# cephfs-data-scan init Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object.
We did not use --force-init and proceeded with scan_extents using a single worker, which was indeed very slow.
After ~24h we interupted the scan_extents and restarted it with 32 workers which went through in about 2h15min w/o any issue.
Then I started scan_inodes with 32 workers this was also finished after ~50min no output on stderr or stdout.
I went on with scan_links, which after ~45 minutes threw the following error:
# cephfs-data-scan scan_links Error ((2) No such file or directory)
Not sure what this indicates necessarily. You can try to get more debug information using: [client] debug mds = 20 debug ms = 1 debug client = 20 in the local ceph.conf for the node running cephfs-data-scan.
then "cephfs-data-scan cleanup" went through w/o any message and took about 9hrs 20min.
Unfortunately, when starting the MDS the cephfs seems still to be in damage. I get quite some "loaded already corrupt dentry:" messages and 2 "[ERR] : bad backtrace on directory inode" errors:
The "corrupt dentry" message is erroneous and fixed already (backports in flight).
(In the following log I removed almost all "loaded already corrupt dentry" entries, for clarity reasons)
2024-06-23T08:06:20.934+0000 7ff05728fb00 0 set uid:gid to 167:167 (ceph:ceph) 2024-06-23T08:06:20.934+0000 7ff05728fb00 0 ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable), process ceph-mds, pid 2 2024-06-23T08:06:20.934+0000 7ff05728fb00 1 main not setting numa affinity 2024-06-23T08:06:20.934+0000 7ff05728fb00 0 pidfile_write: ignore empty --pid-file 2024-06-23T08:06:20.936+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8062 from mon.0 2024-06-23T08:06:21.583+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8063 from mon.0 2024-06-23T08:06:21.583+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Monitors have assigned me to become a standby. 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8064 from mon.0 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:standby --> up:replay 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 replay_start 2024-06-23T08:06:21.604+0000 7ff04bac6700 1 mds.0.8064 waiting for osdmap 34327 (which blocklists prior instance) 2024-06-23T08:06:21.627+0000 7ff0452b9700 0 mds.0.cache creating system inode with ino:0x100 2024-06-23T08:06:21.627+0000 7ff0452b9700 0 mds.0.cache creating system inode with ino:0x1 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.journal EResetJournal 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe start 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe result 2024-06-23T08:06:21.636+0000 7ff0442b7700 1 mds.0.sessionmap wipe done 2024-06-23T08:06:21.656+0000 7ff045aba700 1 mds.0.8064 Finished replaying journal 2024-06-23T08:06:21.656+0000 7ff045aba700 1 mds.0.8064 making mds journal writeable 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8065 from mon.0 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:replay --> up:reconnect 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 reconnect_start 2024-06-23T08:06:22.604+0000 7ff04bac6700 1 mds.0.8064 reopen_log 2024-06-23T08:06:22.605+0000 7ff04bac6700 1 mds.0.8064 reconnect_done 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8066 from mon.0 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map i am now mds.0.8064 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 handle_mds_map state change up:reconnect --> up:rejoin 2024-06-23T08:06:23.605+0000 7ff04bac6700 1 mds.0.8064 rejoin_start 2024-06-23T08:06:23.609+0000 7ff04bac6700 1 mds.0.8064 rejoin_joint_start 2024-06-23T08:06:23.611+0000 7ff045aba700 1 mds.0.cache.den(0x10000000000 groups) loaded already corrupt dentry: [dentry #0x1/data/groups [bf,head] rep@0.0 NULL (dversion lock) pv=0 v= 7910497 ino=(nil) state=0 0x55cf13e9f400] 2024-06-23T08:06:23.611+0000 7ff045aba700 1 mds.0.cache.den(0x1000192ec16.1* scad_prj) loaded already corrupt dentry: [dentry #0x1/home/scad_prj [159,head] rep@0.0 NULL (dversion lock) pv=0 v=2462060 ino=(nil) state=0 0x55cf14220a00] [...] 2024-06-23T08:06:23.628+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 2024-06-23T08:06:23.668+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b [...] 2024-06-23T08:06:23.773+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawn! --- begin dump of recent events --- -9999> 2024-06-23T08:06:23.615+0000 7ff045aba700 1 mds.0.cache.den(0x1000321976d jupyterhub_slurmspawner_67002.log) loaded already corrupt dentry: [dentry #0x1/home/michelotto/jupyterhub_slurmspawner_67002.log [239,36c] rep@0.0 NULL (dversion lock) pv=0 v=26855917 ino=(nil) state=0 0x55cf147ff680] [...] -9878> 2024-06-23T08:06:23.616+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf13f13c00 auth_method 0 [...] -9877> 2024-06-23T08:06:23.616+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf146d1800 auth_method 0 -9458> 2024-06-23T08:06:23.620+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f13400 auth_method 0 -9353> 2024-06-23T08:06:23.622+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf146d0800 auth_method 0 -8980> 2024-06-23T08:06:23.625+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14d9f400 auth_method 0 -8978> 2024-06-23T08:06:23.625+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf14d9fc00 auth_method 0 -8849> 2024-06-23T08:06:23.628+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 -8574> 2024-06-23T08:06:23.633+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14323400 auth_method 0 -8570> 2024-06-23T08:06:23.633+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf1485e400 auth_method 0 -8564> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14322c00 auth_method 0 -8561> 2024-06-23T08:06:23.633+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f12800 auth_method 0 -8555> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14f48000 auth_method 0 -8546> 2024-06-23T08:06:23.633+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf1485f800 auth_method 0 -8541> 2024-06-23T08:06:23.633+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf1516b000 auth_method 0 -8470> 2024-06-23T08:06:23.634+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf14322400 auth_method 0 -8451> 2024-06-23T08:06:23.635+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf13f13800 auth_method 0 -8445> 2024-06-23T08:06:23.635+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf1485e800 auth_method 0 -8243> 2024-06-23T08:06:23.637+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf141ce400 auth_method 0 -7381> 2024-06-23T08:06:23.645+0000 7ff04f2cd700 10 monclient: get_auth_request con 0x55cf1485f400 auth_method 0 -6319> 2024-06-23T08:06:23.660+0000 7ff04eacc700 10 monclient: get_auth_request con 0x55cf14f48400 auth_method 0 -5946> 2024-06-23T08:06:23.666+0000 7ff04e2cb700 10 monclient: get_auth_request con 0x55cf146d1400 auth_method 0 [...] -5677> 2024-06-23T08:06:23.668+0000 7ff045aba700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b [...] -8> 2024-06-23T08:06:23.753+0000 7ff045aba700 5 mds.beacon.default.cephmon-01.cepqjp set_want_state: up:rejoin -> down:damaged -7> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client log_queue is 2 last_log 2 sent 0 num 2 unsent 2 sending 2 -6> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client will send 2024-06-23T08:06:23.629743+0000 mds.default.cephmon-01.cepqjp (mds.0) 1 : cluster [ERR] bad backtrace on directory inode 0x10003e42340 -5> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 log_client will send 2024-06-23T08:06:23.669673+0000 mds.default.cephmon-01.cepqjp (mds.0) 2 : cluster [ERR] bad backtrace on directory inode 0x10003e45d8b -4> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 monclient: _send_mon_message to mon.cephmon-01 at v2:10.1.3.21:3300/0 -3> 2024-06-23T08:06:23.753+0000 7ff045aba700 5 mds.beacon.default.cephmon-01.cepqjp Sending beacon down:damaged seq 4 -2> 2024-06-23T08:06:23.753+0000 7ff045aba700 10 monclient: _send_mon_message to mon.cephmon-01 at v2:10.1.3.21:3300/0 -1> 2024-06-23T08:06:23.773+0000 7ff04e2cb700 5 mds.beacon.default.cephmon-01.cepqjp received beacon reply down:damaged seq 4 rtt 0.0200001 0> 2024-06-23T08:06:23.773+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawn! --- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7ff045aba700 / MR_Finisher 7ff04e2cb700 / msgr-worker-2 7ff04eacc700 / msgr-worker-1 7ff04f2cd700 / msgr-worker-0 max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-01.cepqjp.log --- end dump of recent events --- 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp e: '/usr/bin/ceph-mds' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 0: '/usr/bin/ceph-mds' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 1: '-n' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 2: 'mds.default.cephmon-01.cepqjp' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 3: '-f' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 4: '--setuser' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 5: 'ceph' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 6: '--setgroup' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 7: 'ceph' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 8: '--default-log-to-file=false' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 9: '--default-log-to-journald=true' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp 10: '--default-log-to-stderr=false' 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp respawning with exe /usr/bin/ceph-mds 2024-06-23T08:06:23.786+0000 7ff045aba700 1 mds.default.cephmon-01.cepqjp exe_path /proc/self/exe 2024-06-23T08:06:23.812+0000 7fe58d619b00 0 ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable), process ceph-mds, pid 2 2024-06-23T08:06:23.812+0000 7fe58d619b00 1 main not setting numa affinity 2024-06-23T08:06:23.813+0000 7fe58d619b00 0 pidfile_write: ignore empty --pid-file 2024-06-23T08:06:23.814+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8067 from mon.0 2024-06-23T08:06:24.772+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Updating MDS map to version 8068 from mon.0 2024-06-23T08:06:24.772+0000 7fe58226e700 1 mds.default.cephmon-01.cepqjp Monitors have assigned me to become a standby. 2024-06-23T08:49:28.778+0000 7fe584272700 1 mds.default.cephmon-01.cepqjp asok_command: heap {heapcmd=stats,prefix=heap} (starting...) 2024-06-23T22:00:04.664+0000 7fe583a71700 -1 received signal: Hangup from Kernel ( Could be generated by pthread_kill(), raise(), abort(), alarm() ) UID: 0
Any ideas how to proceed?
Whatever you snipped from the log has the real error. The MDS tries to recover from both "ERR" messages you are concerned about. It should not go damaged. -- Patrick Donnelly, Ph.D. He / Him / His Red Hat Partner Engineer IBM, Inc. GPG: 19F28A586F808C2402351B93C3301A3E258DD79D
Hi Patrick, Xiubo and List, finally we managed to get the filesystem repaired and running again! YEAH, I'm so happy!! Big thanks for your support Patrick and Xiubo! (Would love invite you for a beer)! Please see some comments and (important?) questions below: On 6/25/24 03:14, Patrick Donnelly wrote:
On Mon, Jun 24, 2024 at 5:22 PM Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
(resending this, the original message seems that it didn't make it through between all the SPAM recently sent to the list, my apologies if it doubles at some point)
Hi List,
we are still struggeling to get our cephfs back online again, this is an update to inform you what we did so far, and we kindly ask for any input on this to get an idea on how to proceed:
After resetting the journals Xiubo suggested (in a PM) to go on with the disaster recovery procedure:
cephfs-data-scan init skipped creating the inodes 0x0x1 and 0x0x100
[root@ceph01-b ~]# cephfs-data-scan init Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object.
We did not use --force-init and proceeded with scan_extents using a single worker, which was indeed very slow.
After ~24h we interupted the scan_extents and restarted it with 32 workers which went through in about 2h15min w/o any issue.
Then I started scan_inodes with 32 workers this was also finished after ~50min no output on stderr or stdout.
I went on with scan_links, which after ~45 minutes threw the following error:
# cephfs-data-scan scan_links Error ((2) No such file or directory)
Not sure what this indicates necessarily. You can try to get more debug information using:
[client] debug mds = 20 debug ms = 1 debug client = 20
in the local ceph.conf for the node running cephfs-data-scan.
I did that, and restarted the "cephfs-data-scan scan_links" . It didn't produce any additional debug output, however this time it just went through without error (~50 min) We then reran "cephfs-data-scan cleanup" and it also finished without error after about 10h. We then set the fs as repaired and all seems to work fin again: [root@ceph01-b ~]# ceph mds repaired 0 repaired: restoring rank 1:0 [root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_OK services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 6d) mgr: cephmon-01.dsxcho(active, since 6d), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 1/1 daemons up, 5 standby osd: 336 osds: 336 up (since 2M), 336 in (since 4M) data: volumes: 1/1 healthy pools: 4 pools, 6401 pgs objects: 284.68M objects, 623 TiB usage: 890 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6206 active+clean 140 active+clean+scrubbing 55 active+clean+scrubbing+deep io: client: 3.9 MiB/s rd, 84 B/s wr, 482 op/s rd, 1.11k op/s wr [root@ceph01-b ~]# ceph fs status cephfs - 0 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 active default.cephmon-03.xcujhz Reqs: 0 /s 124k 60.3k 1993 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 298G 63.4T sdd-rep-data-pool data 10.2T 84.5T hdd-ec-data-pool data 808T 1929T STANDBY MDS default.cephmon-01.cepqjp default.cephmon-01.pvnqad default.cephmon-02.duujba default.cephmon-02.nyfook default.cephmon-03.chjusj MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) The msd log however shows some "bad backtrace on directory inode" messages: 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8082 from mon.1 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:standby --> up:replay 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 replay_start 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 waiting for osdmap 34331 (which blocklists prior instance) 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x100 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x1 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.journal EResetJournal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe start 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe result 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe done 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 Finished replaying journal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 making mds journal writeable 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8083 from mon.1 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:replay --> up:reconnect 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_start 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reopen_log 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_done 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8084 from mon.1 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:reconnect --> up:rejoin 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 rejoin_start 2024-06-25T18:45:38.583+0000 7f8594659700 1 mds.0.8082 rejoin_joint_start 2024-06-25T18:45:38.592+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 2024-06-25T18:45:38.680+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d90 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d9f 2024-06-25T18:45:38.785+0000 7f858fe50700 1 mds.0.8082 rejoin_done 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8085 from mon.1 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:rejoin --> up:active 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 recovery_done -- successful recovery! 2024-06-25T18:45:39.584+0000 7f8594659700 1 mds.0.8082 active_start 2024-06-25T18:45:39.585+0000 7f8594659700 1 mds.0.8082 cluster recovered. 2024-06-25T18:45:42.409+0000 7f8591e54700 -1 mds.pinger is_rank_lagging: rank=0 was never sent ping request. 2024-06-25T18:57:28.213+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x4 Is there anything that we can do about this, to get rid of the "bad backtrace on directory inode"? Sone more question: 1. As Xiubo suggested, we now tried to mount the filesystem with the "nowsysnc" option <https://tracker.ceph.com/issues/61009#note-26>: [root@ceph01-b ~]# mount -t ceph cephfs_user@.cephfs=/ /mnt/cephfs -o secretfile=/etc/ceph/ceph.client.cephfs_user.secret,nowsync however the option seems not to show up in /proc/mounts [root@ceph01-b ~]# grep ceph /proc/mounts cephfs_user@aae23c5c-a98b-11ee-b44d-00620b05cac4.cephfs=/ /mnt/cephfs ceph rw,relatime,name=cephfs_user,secret=<hidden>,ms_mode=prefer-crc,acl,mon_addr=10.1.3.21:3300/10.1.3.22:3300/10.1.3.23:3300 0 0 The kernel version is 5.14.0 (from Rocky 9.3) [root@ceph01-b ~]# uname -a Linux ceph01-b 5.14.0-362.24.1.el9_3.x86_64 #1 SMP PREEMPT_DYNAMIC Wed Mar 13 17:33:16 UTC 2024 x86_64 x86_64 x86_64 GNU/Linux Is this expected? How can we make sure that the filesystem uses 'nowsync', so that we do not hit the bug <https://tracker.ceph.com/issues/61009> again? 2. There are two empty files in lost+found now. Is ist save to remove them? [root@ceph01-b lost+found]# ls -la total 0 drwxr-xr-x 2 root root 1 Jan 1 1970 . drwxr-xr-x 4 root root 2 Mar 13 21:22 .. -r-x------ 1 root root 0 Jun 20 23:50 100037a50e2 -r-x------ 1 root root 0 Jun 20 19:05 200049612e5 3. Are there any specific steps that we should perform now (scrub or similar things) before we put the filesystem into production again? Best & thanks again Dietmar
...sending also to the list and Xiubo (were accidentally removed from recipients)... On 6/25/24 21:28, Dietmar Rieder wrote:
Hi Patrick, Xiubo and List,
finally we managed to get the filesystem repaired and running again! YEAH, I'm so happy!!
Big thanks for your support Patrick and Xiubo! (Would love invite you for a beer)!
Please see some comments and (important?) questions below:
On 6/25/24 03:14, Patrick Donnelly wrote:
On Mon, Jun 24, 2024 at 5:22 PM Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
(resending this, the original message seems that it didn't make it through between all the SPAM recently sent to the list, my apologies if it doubles at some point)
Hi List,
we are still struggeling to get our cephfs back online again, this is an update to inform you what we did so far, and we kindly ask for any input on this to get an idea on how to proceed:
After resetting the journals Xiubo suggested (in a PM) to go on with the disaster recovery procedure:
cephfs-data-scan init skipped creating the inodes 0x0x1 and 0x0x100
[root@ceph01-b ~]# cephfs-data-scan init Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object.
We did not use --force-init and proceeded with scan_extents using a single worker, which was indeed very slow.
After ~24h we interupted the scan_extents and restarted it with 32 workers which went through in about 2h15min w/o any issue.
Then I started scan_inodes with 32 workers this was also finished after ~50min no output on stderr or stdout.
I went on with scan_links, which after ~45 minutes threw the following error:
# cephfs-data-scan scan_links Error ((2) No such file or directory)
Not sure what this indicates necessarily. You can try to get more debug information using:
[client] debug mds = 20 debug ms = 1 debug client = 20
in the local ceph.conf for the node running cephfs-data-scan.
I did that, and restarted the "cephfs-data-scan scan_links" .
It didn't produce any additional debug output, however this time it just went through without error (~50 min)
We then reran "cephfs-data-scan cleanup" and it also finished without error after about 10h.
We then set the fs as repaired and all seems to work fin again:
[root@ceph01-b ~]# ceph mds repaired 0 repaired: restoring rank 1:0
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_OK
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 6d) mgr: cephmon-01.dsxcho(active, since 6d), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 1/1 daemons up, 5 standby osd: 336 osds: 336 up (since 2M), 336 in (since 4M)
data: volumes: 1/1 healthy pools: 4 pools, 6401 pgs objects: 284.68M objects, 623 TiB usage: 890 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6206 active+clean 140 active+clean+scrubbing 55 active+clean+scrubbing+deep
io: client: 3.9 MiB/s rd, 84 B/s wr, 482 op/s rd, 1.11k op/s wr
[root@ceph01-b ~]# ceph fs status cephfs - 0 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 active default.cephmon-03.xcujhz Reqs: 0 /s 124k 60.3k 1993 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 298G 63.4T sdd-rep-data-pool data 10.2T 84.5T hdd-ec-data-pool data 808T 1929T STANDBY MDS default.cephmon-01.cepqjp default.cephmon-01.pvnqad default.cephmon-02.duujba default.cephmon-02.nyfook default.cephmon-03.chjusj MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The msd log however shows some "bad backtrace on directory inode" messages:
2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8082 from mon.1 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:standby --> up:replay 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 replay_start 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 waiting for osdmap 34331 (which blocklists prior instance) 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x100 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x1 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.journal EResetJournal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe start 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe result 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe done 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 Finished replaying journal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 making mds journal writeable 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8083 from mon.1 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:replay --> up:reconnect 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_start 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reopen_log 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_done 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8084 from mon.1 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:reconnect --> up:rejoin 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 rejoin_start 2024-06-25T18:45:38.583+0000 7f8594659700 1 mds.0.8082 rejoin_joint_start 2024-06-25T18:45:38.592+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 2024-06-25T18:45:38.680+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d90 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d9f 2024-06-25T18:45:38.785+0000 7f858fe50700 1 mds.0.8082 rejoin_done 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8085 from mon.1 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:rejoin --> up:active 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 recovery_done -- successful recovery! 2024-06-25T18:45:39.584+0000 7f8594659700 1 mds.0.8082 active_start 2024-06-25T18:45:39.585+0000 7f8594659700 1 mds.0.8082 cluster recovered. 2024-06-25T18:45:42.409+0000 7f8591e54700 -1 mds.pinger is_rank_lagging: rank=0 was never sent ping request. 2024-06-25T18:57:28.213+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x4
Is there anything that we can do about this, to get rid of the "bad backtrace on directory inode"?
Sone more question:
1. As Xiubo suggested, we now tried to mount the filesystem with the "nowsysnc" option <https://tracker.ceph.com/issues/61009#note-26>:
[root@ceph01-b ~]# mount -t ceph cephfs_user@.cephfs=/ /mnt/cephfs -o secretfile=/etc/ceph/ceph.client.cephfs_user.secret,nowsync
however the option seems not to show up in /proc/mounts
[root@ceph01-b ~]# grep ceph /proc/mounts cephfs_user@aae23c5c-a98b-11ee-b44d-00620b05cac4.cephfs=/ /mnt/cephfs ceph rw,relatime,name=cephfs_user,secret=<hidden>,ms_mode=prefer-crc,acl,mon_addr=10.1.3.21:3300/10.1.3.22:3300/10.1.3.23:3300 0 0
The kernel version is 5.14.0 (from Rocky 9.3)
[root@ceph01-b ~]# uname -a Linux ceph01-b 5.14.0-362.24.1.el9_3.x86_64 #1 SMP PREEMPT_DYNAMIC Wed Mar 13 17:33:16 UTC 2024 x86_64 x86_64 x86_64 GNU/Linux
Is this expected? How can we make sure that the filesystem uses 'nowsync', so that we do not hit the bug <https://tracker.ceph.com/issues/61009> again?
Oh, I think I misunderstood the suggested workaround. I guess we need to disable "nowsync", which is set by default, right? so: -o wsync should be the workaround, right?
2. There are two empty files in lost+found now. Is ist save to remove them?
[root@ceph01-b lost+found]# ls -la total 0 drwxr-xr-x 2 root root 1 Jan 1 1970 . drwxr-xr-x 4 root root 2 Mar 13 21:22 .. -r-x------ 1 root root 0 Jun 20 23:50 100037a50e2 -r-x------ 1 root root 0 Jun 20 19:05 200049612e5
3. Are there any specific steps that we should perform now (scrub or similar things) before we put the filesystem into production again?
Dietmar
Can anybody comment on my questions below? Thanks so much in advance.... Am 26. Juni 2024 08:08:39 MESZ schrieb Dietmar Rieder <dietmar.rieder@i-med.ac.at>:
...sending also to the list and Xiubo (were accidentally removed from recipients)...
On 6/25/24 21:28, Dietmar Rieder wrote:
Hi Patrick, Xiubo and List,
finally we managed to get the filesystem repaired and running again! YEAH, I'm so happy!!
Big thanks for your support Patrick and Xiubo! (Would love invite you for a beer)!
Please see some comments and (important?) questions below:
On 6/25/24 03:14, Patrick Donnelly wrote:
On Mon, Jun 24, 2024 at 5:22 PM Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
(resending this, the original message seems that it didn't make it through between all the SPAM recently sent to the list, my apologies if it doubles at some point)
Hi List,
we are still struggeling to get our cephfs back online again, this is an update to inform you what we did so far, and we kindly ask for any input on this to get an idea on how to proceed:
After resetting the journals Xiubo suggested (in a PM) to go on with the disaster recovery procedure:
cephfs-data-scan init skipped creating the inodes 0x0x1 and 0x0x100
[root@ceph01-b ~]# cephfs-data-scan init Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object.
We did not use --force-init and proceeded with scan_extents using a single worker, which was indeed very slow.
After ~24h we interupted the scan_extents and restarted it with 32 workers which went through in about 2h15min w/o any issue.
Then I started scan_inodes with 32 workers this was also finished after ~50min no output on stderr or stdout.
I went on with scan_links, which after ~45 minutes threw the following error:
# cephfs-data-scan scan_links Error ((2) No such file or directory)
Not sure what this indicates necessarily. You can try to get more debug information using:
[client] debug mds = 20 debug ms = 1 debug client = 20
in the local ceph.conf for the node running cephfs-data-scan.
I did that, and restarted the "cephfs-data-scan scan_links" .
It didn't produce any additional debug output, however this time it just went through without error (~50 min)
We then reran "cephfs-data-scan cleanup" and it also finished without error after about 10h.
We then set the fs as repaired and all seems to work fin again:
[root@ceph01-b ~]# ceph mds repaired 0 repaired: restoring rank 1:0
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_OK
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 6d) mgr: cephmon-01.dsxcho(active, since 6d), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 1/1 daemons up, 5 standby osd: 336 osds: 336 up (since 2M), 336 in (since 4M)
data: volumes: 1/1 healthy pools: 4 pools, 6401 pgs objects: 284.68M objects, 623 TiB usage: 890 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6206 active+clean 140 active+clean+scrubbing 55 active+clean+scrubbing+deep
io: client: 3.9 MiB/s rd, 84 B/s wr, 482 op/s rd, 1.11k op/s wr
[root@ceph01-b ~]# ceph fs status cephfs - 0 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 active default.cephmon-03.xcujhz Reqs: 0 /s 124k 60.3k 1993 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 298G 63.4T sdd-rep-data-pool data 10.2T 84.5T hdd-ec-data-pool data 808T 1929T STANDBY MDS default.cephmon-01.cepqjp default.cephmon-01.pvnqad default.cephmon-02.duujba default.cephmon-02.nyfook default.cephmon-03.chjusj MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The msd log however shows some "bad backtrace on directory inode" messages:
2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8082 from mon.1 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:standby --> up:replay 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 replay_start 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 waiting for osdmap 34331 (which blocklists prior instance) 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x100 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x1 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.journal EResetJournal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe start 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe result 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe done 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 Finished replaying journal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 making mds journal writeable 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8083 from mon.1 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:replay --> up:reconnect 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_start 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reopen_log 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_done 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8084 from mon.1 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:reconnect --> up:rejoin 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 rejoin_start 2024-06-25T18:45:38.583+0000 7f8594659700 1 mds.0.8082 rejoin_joint_start 2024-06-25T18:45:38.592+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 2024-06-25T18:45:38.680+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d90 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d9f 2024-06-25T18:45:38.785+0000 7f858fe50700 1 mds.0.8082 rejoin_done 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8085 from mon.1 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:rejoin --> up:active 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 recovery_done -- successful recovery! 2024-06-25T18:45:39.584+0000 7f8594659700 1 mds.0.8082 active_start 2024-06-25T18:45:39.585+0000 7f8594659700 1 mds.0.8082 cluster recovered. 2024-06-25T18:45:42.409+0000 7f8591e54700 -1 mds.pinger is_rank_lagging: rank=0 was never sent ping request. 2024-06-25T18:57:28.213+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x4
Is there anything that we can do about this, to get rid of the "bad backtrace on directory inode"?
Sone more question:
1. As Xiubo suggested, we now tried to mount the filesystem with the "nowsysnc" option <https://tracker.ceph.com/issues/61009#note-26>:
[root@ceph01-b ~]# mount -t ceph cephfs_user@.cephfs=/ /mnt/cephfs -o secretfile=/etc/ceph/ceph.client.cephfs_user.secret,nowsync
however the option seems not to show up in /proc/mounts
[root@ceph01-b ~]# grep ceph /proc/mounts cephfs_user@aae23c5c-a98b-11ee-b44d-00620b05cac4.cephfs=/ /mnt/cephfs ceph rw,relatime,name=cephfs_user,secret=<hidden>,ms_mode=prefer-crc,acl,mon_addr=10.1.3.21:3300/10.1.3.22:3300/10.1.3.23:3300 0 0
The kernel version is 5.14.0 (from Rocky 9.3)
[root@ceph01-b ~]# uname -a Linux ceph01-b 5.14.0-362.24.1.el9_3.x86_64 #1 SMP PREEMPT_DYNAMIC Wed Mar 13 17:33:16 UTC 2024 x86_64 x86_64 x86_64 GNU/Linux
Is this expected? How can we make sure that the filesystem uses 'nowsync', so that we do not hit the bug <https://tracker.ceph.com/issues/61009> again?
Oh, I think I misunderstood the suggested workaround. I guess we need to disable "nowsync", which is set by default, right?
so: -o wsync
should be the workaround, right?
2. There are two empty files in lost+found now. Is ist save to remove them?
[root@ceph01-b lost+found]# ls -la total 0 drwxr-xr-x 2 root root 1 Jan 1 1970 . drwxr-xr-x 4 root root 2 Mar 13 21:22 .. -r-x------ 1 root root 0 Jun 20 23:50 100037a50e2 -r-x------ 1 root root 0 Jun 20 19:05 200049612e5
3. Are there any specific steps that we should perform now (scrub or similar things) before we put the filesystem into production again?
Dietmar
Hi Dietmar, I understand the option to be set is 'wsync', not 'nowsync'. See https://docs.ceph.com/en/latest/man/8/mount.ceph/ nowsync enables async dirops, which is what triggers the assertion in https://tracker.ceph.com/issues/61009 The reason why you don't see it in /proc/mounts is because it is the default in recent kernels (see https://github.com/gregkh/linux/commit/f7a67b463fb83a4b9b11ceaa8ec4950b8fb7f...) If you set 'wsync' among your mount options, this will show up in /proc/mounts Cheers, Enrico On 6/27/24 06:37, Dietmar Rieder wrote:
Can anybody comment on my questions below? Thanks so much in advance....
Am 26. Juni 2024 08:08:39 MESZ schrieb Dietmar Rieder <dietmar.rieder@i-med.ac.at>:
...sending also to the list and Xiubo (were accidentally removed from recipients)...
On 6/25/24 21:28, Dietmar Rieder wrote:
Hi Patrick, Xiubo and List,
finally we managed to get the filesystem repaired and running again! YEAH, I'm so happy!!
Big thanks for your support Patrick and Xiubo! (Would love invite you for a beer)!
Please see some comments and (important?) questions below:
On 6/25/24 03:14, Patrick Donnelly wrote:
On Mon, Jun 24, 2024 at 5:22 PM Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
(resending this, the original message seems that it didn't make it through between all the SPAM recently sent to the list, my apologies if it doubles at some point)
Hi List,
we are still struggeling to get our cephfs back online again, this is an update to inform you what we did so far, and we kindly ask for any input on this to get an idea on how to proceed:
After resetting the journals Xiubo suggested (in a PM) to go on with the disaster recovery procedure:
cephfs-data-scan init skipped creating the inodes 0x0x1 and 0x0x100
[root@ceph01-b ~]# cephfs-data-scan init Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object.
We did not use --force-init and proceeded with scan_extents using a single worker, which was indeed very slow.
After ~24h we interupted the scan_extents and restarted it with 32 workers which went through in about 2h15min w/o any issue.
Then I started scan_inodes with 32 workers this was also finished after ~50min no output on stderr or stdout.
I went on with scan_links, which after ~45 minutes threw the following error:
# cephfs-data-scan scan_links Error ((2) No such file or directory) Not sure what this indicates necessarily. You can try to get more debug information using:
[client] debug mds = 20 debug ms = 1 debug client = 20
in the local ceph.conf for the node running cephfs-data-scan. I did that, and restarted the "cephfs-data-scan scan_links" .
It didn't produce any additional debug output, however this time it just went through without error (~50 min)
We then reran "cephfs-data-scan cleanup" and it also finished without error after about 10h.
We then set the fs as repaired and all seems to work fin again:
[root@ceph01-b ~]# ceph mds repaired 0 repaired: restoring rank 1:0
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_OK
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 6d) mgr: cephmon-01.dsxcho(active, since 6d), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 1/1 daemons up, 5 standby osd: 336 osds: 336 up (since 2M), 336 in (since 4M)
data: volumes: 1/1 healthy pools: 4 pools, 6401 pgs objects: 284.68M objects, 623 TiB usage: 890 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6206 active+clean 140 active+clean+scrubbing 55 active+clean+scrubbing+deep
io: client: 3.9 MiB/s rd, 84 B/s wr, 482 op/s rd, 1.11k op/s wr
[root@ceph01-b ~]# ceph fs status cephfs - 0 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 active default.cephmon-03.xcujhz Reqs: 0 /s 124k 60.3k 1993 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 298G 63.4T sdd-rep-data-pool data 10.2T 84.5T hdd-ec-data-pool data 808T 1929T STANDBY MDS default.cephmon-01.cepqjp default.cephmon-01.pvnqad default.cephmon-02.duujba default.cephmon-02.nyfook default.cephmon-03.chjusj MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The msd log however shows some "bad backtrace on directory inode" messages:
2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8082 from mon.1 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:standby --> up:replay 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 replay_start 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 waiting for osdmap 34331 (which blocklists prior instance) 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x100 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x1 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.journal EResetJournal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe start 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe result 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe done 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 Finished replaying journal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 making mds journal writeable 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8083 from mon.1 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:replay --> up:reconnect 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_start 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reopen_log 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_done 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8084 from mon.1 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:reconnect --> up:rejoin 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 rejoin_start 2024-06-25T18:45:38.583+0000 7f8594659700 1 mds.0.8082 rejoin_joint_start 2024-06-25T18:45:38.592+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 2024-06-25T18:45:38.680+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d90 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d9f 2024-06-25T18:45:38.785+0000 7f858fe50700 1 mds.0.8082 rejoin_done 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8085 from mon.1 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:rejoin --> up:active 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 recovery_done -- successful recovery! 2024-06-25T18:45:39.584+0000 7f8594659700 1 mds.0.8082 active_start 2024-06-25T18:45:39.585+0000 7f8594659700 1 mds.0.8082 cluster recovered. 2024-06-25T18:45:42.409+0000 7f8591e54700 -1 mds.pinger is_rank_lagging: rank=0 was never sent ping request. 2024-06-25T18:57:28.213+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x4
Is there anything that we can do about this, to get rid of the "bad backtrace on directory inode"?
Sone more question:
1. As Xiubo suggested, we now tried to mount the filesystem with the "nowsysnc" option <https://tracker.ceph.com/issues/61009#note-26>:
[root@ceph01-b ~]# mount -t ceph cephfs_user@.cephfs=/ /mnt/cephfs -o secretfile=/etc/ceph/ceph.client.cephfs_user.secret,nowsync
however the option seems not to show up in /proc/mounts
[root@ceph01-b ~]# grep ceph /proc/mounts cephfs_user@aae23c5c-a98b-11ee-b44d-00620b05cac4.cephfs=/ /mnt/cephfs ceph rw,relatime,name=cephfs_user,secret=<hidden>,ms_mode=prefer-crc,acl,mon_addr=10.1.3.21:3300/10.1.3.22:3300/10.1.3.23:3300 0 0
The kernel version is 5.14.0 (from Rocky 9.3)
[root@ceph01-b ~]# uname -a Linux ceph01-b 5.14.0-362.24.1.el9_3.x86_64 #1 SMP PREEMPT_DYNAMIC Wed Mar 13 17:33:16 UTC 2024 x86_64 x86_64 x86_64 GNU/Linux
Is this expected? How can we make sure that the filesystem uses 'nowsync', so that we do not hit the bug <https://tracker.ceph.com/issues/61009> again?
Oh, I think I misunderstood the suggested workaround. I guess we need to disable "nowsync", which is set by default, right?
so: -o wsync
should be the workaround, right?
2. There are two empty files in lost+found now. Is ist save to remove them?
[root@ceph01-b lost+found]# ls -la total 0 drwxr-xr-x 2 root root 1 Jan 1 1970 . drwxr-xr-x 4 root root 2 Mar 13 21:22 .. -r-x------ 1 root root 0 Jun 20 23:50 100037a50e2 -r-x------ 1 root root 0 Jun 20 19:05 200049612e5
3. Are there any specific steps that we should perform now (scrub or similar things) before we put the filesystem into production again?
Dietmar
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Enrico Bocchi CERN European Laboratory for Particle Physics IT - Storage & Data Management - General Storage Services Mailbox: G20500 - Office: 31-2-010 1211 Genève 23 Switzerland
Hi Enrico, thanks so much for your comment. You are right, that's what I figured out a bit later, see below. BTW, I was able to repair the filesystem and all is working fine again, it seems that we did not lose any data (will post a summary, for the record) Thanks again, DIetmar On 6/28/24 13:08, Enrico Bocchi wrote:
Hi Dietmar,
I understand the option to be set is 'wsync', not 'nowsync'. See https://docs.ceph.com/en/latest/man/8/mount.ceph/ nowsync enables async dirops, which is what triggers the assertion in https://tracker.ceph.com/issues/61009
The reason why you don't see it in /proc/mounts is because it is the default in recent kernels (see https://github.com/gregkh/linux/commit/f7a67b463fb83a4b9b11ceaa8ec4950b8fb7f...) If you set 'wsync' among your mount options, this will show up in /proc/mounts
Cheers, Enrico
On 6/27/24 06:37, Dietmar Rieder wrote:
[...]
Oh, I think I misunderstood the suggested workaround. I guess we need to disable "nowsync", which is set by default, right?
so: -o wsync
should be the workaround, right?
On 6/26/24 14:08, Dietmar Rieder wrote:
...sending also to the list and Xiubo (were accidentally removed from recipients)...
On 6/25/24 21:28, Dietmar Rieder wrote:
Hi Patrick, Xiubo and List,
finally we managed to get the filesystem repaired and running again! YEAH, I'm so happy!!
Big thanks for your support Patrick and Xiubo! (Would love invite you for a beer)!
Please see some comments and (important?) questions below:
On 6/25/24 03:14, Patrick Donnelly wrote:
On Mon, Jun 24, 2024 at 5:22 PM Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
(resending this, the original message seems that it didn't make it through between all the SPAM recently sent to the list, my apologies if it doubles at some point)
Hi List,
we are still struggeling to get our cephfs back online again, this is an update to inform you what we did so far, and we kindly ask for any input on this to get an idea on how to proceed:
After resetting the journals Xiubo suggested (in a PM) to go on with the disaster recovery procedure:
cephfs-data-scan init skipped creating the inodes 0x0x1 and 0x0x100
[root@ceph01-b ~]# cephfs-data-scan init Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object.
We did not use --force-init and proceeded with scan_extents using a single worker, which was indeed very slow.
After ~24h we interupted the scan_extents and restarted it with 32 workers which went through in about 2h15min w/o any issue.
Then I started scan_inodes with 32 workers this was also finished after ~50min no output on stderr or stdout.
I went on with scan_links, which after ~45 minutes threw the following error:
# cephfs-data-scan scan_links Error ((2) No such file or directory)
Not sure what this indicates necessarily. You can try to get more debug information using:
[client] debug mds = 20 debug ms = 1 debug client = 20
in the local ceph.conf for the node running cephfs-data-scan.
I did that, and restarted the "cephfs-data-scan scan_links" .
It didn't produce any additional debug output, however this time it just went through without error (~50 min)
We then reran "cephfs-data-scan cleanup" and it also finished without error after about 10h.
We then set the fs as repaired and all seems to work fin again:
[root@ceph01-b ~]# ceph mds repaired 0 repaired: restoring rank 1:0
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_OK
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 6d) mgr: cephmon-01.dsxcho(active, since 6d), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 1/1 daemons up, 5 standby osd: 336 osds: 336 up (since 2M), 336 in (since 4M)
data: volumes: 1/1 healthy pools: 4 pools, 6401 pgs objects: 284.68M objects, 623 TiB usage: 890 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6206 active+clean 140 active+clean+scrubbing 55 active+clean+scrubbing+deep
io: client: 3.9 MiB/s rd, 84 B/s wr, 482 op/s rd, 1.11k op/s wr
[root@ceph01-b ~]# ceph fs status cephfs - 0 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 active default.cephmon-03.xcujhz Reqs: 0 /s 124k 60.3k 1993 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 298G 63.4T sdd-rep-data-pool data 10.2T 84.5T hdd-ec-data-pool data 808T 1929T STANDBY MDS default.cephmon-01.cepqjp default.cephmon-01.pvnqad default.cephmon-02.duujba default.cephmon-02.nyfook default.cephmon-03.chjusj MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The msd log however shows some "bad backtrace on directory inode" messages:
2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8082 from mon.1 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:standby --> up:replay 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 replay_start 2024-06-25T18:45:36.575+0000 7f8594659700 1 mds.0.8082 waiting for osdmap 34331 (which blocklists prior instance) 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x100 2024-06-25T18:45:36.581+0000 7f858de4c700 0 mds.0.cache creating system inode with ino:0x1 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.journal EResetJournal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe start 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe result 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.sessionmap wipe done 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 Finished replaying journal 2024-06-25T18:45:36.589+0000 7f858ce4a700 1 mds.0.8082 making mds journal writeable 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8083 from mon.1 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:replay --> up:reconnect 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_start 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reopen_log 2024-06-25T18:45:37.578+0000 7f8594659700 1 mds.0.8082 reconnect_done 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8084 from mon.1 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:reconnect --> up:rejoin 2024-06-25T18:45:38.579+0000 7f8594659700 1 mds.0.8082 rejoin_start 2024-06-25T18:45:38.583+0000 7f8594659700 1 mds.0.8082 rejoin_joint_start 2024-06-25T18:45:38.592+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 2024-06-25T18:45:38.680+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d90 2024-06-25T18:45:38.754+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d9f 2024-06-25T18:45:38.785+0000 7f858fe50700 1 mds.0.8082 rejoin_done 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.default.cephmon-03.xcujhz Updating MDS map to version 8085 from mon.1 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map i am now mds.0.8082 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 handle_mds_map state change up:rejoin --> up:active 2024-06-25T18:45:39.582+0000 7f8594659700 1 mds.0.8082 recovery_done -- successful recovery! 2024-06-25T18:45:39.584+0000 7f8594659700 1 mds.0.8082 active_start 2024-06-25T18:45:39.585+0000 7f8594659700 1 mds.0.8082 cluster recovered. 2024-06-25T18:45:42.409+0000 7f8591e54700 -1 mds.pinger is_rank_lagging: rank=0 was never sent ping request. 2024-06-25T18:57:28.213+0000 7f858e64d700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x4
Is there anything that we can do about this, to get rid of the "bad backtrace on directory inode"?
Sone more question:
1. As Xiubo suggested, we now tried to mount the filesystem with the "nowsysnc" option <https://tracker.ceph.com/issues/61009#note-26>:
[root@ceph01-b ~]# mount -t ceph cephfs_user@.cephfs=/ /mnt/cephfs -o secretfile=/etc/ceph/ceph.client.cephfs_user.secret,nowsync
however the option seems not to show up in /proc/mounts
[root@ceph01-b ~]# grep ceph /proc/mounts cephfs_user@aae23c5c-a98b-11ee-b44d-00620b05cac4.cephfs=/ /mnt/cephfs ceph rw,relatime,name=cephfs_user,secret=<hidden>,ms_mode=prefer-crc,acl,mon_addr=10.1.3.21:3300/10.1.3.22:3300/10.1.3.23:3300 0 0
The kernel version is 5.14.0 (from Rocky 9.3)
[root@ceph01-b ~]# uname -a Linux ceph01-b 5.14.0-362.24.1.el9_3.x86_64 #1 SMP PREEMPT_DYNAMIC Wed Mar 13 17:33:16 UTC 2024 x86_64 x86_64 x86_64 GNU/Linux
Is this expected? How can we make sure that the filesystem uses 'nowsync', so that we do not hit the bug <https://tracker.ceph.com/issues/61009> again?
Oh, I think I misunderstood the suggested workaround. I guess we need to disable "nowsync", which is set by default, right?
so: -o wsync
should be the workaround, right?
Yeah, right.
2. There are two empty files in lost+found now. Is ist save to remove them?
[root@ceph01-b lost+found]# ls -la total 0 drwxr-xr-x 2 root root 1 Jan 1 1970 . drwxr-xr-x 4 root root 2 Mar 13 21:22 .. -r-x------ 1 root root 0 Jun 20 23:50 100037a50e2 -r-x------ 1 root root 0 Jun 20 19:05 200049612e5
3. Are there any specific steps that we should perform now (scrub or similar things) before we put the filesystem into production again?
Dietmar
Hi Dietmar, have you already blocked all cephfs clients? Joachim *Joachim Kraftmayer* CEO | p: +49 89 2152527-21 | e: joachim.kraftmayer@clyso.com a: Loristr. 8 | 80335 Munich | Germany | w: https://clyso.com | Utting a. A. | HR: Augsburg | HRB 25866 | USt. ID: DE275430677 Am Mi., 19. Juni 2024 um 09:44 Uhr schrieb Dietmar Rieder < dietmar.rieder@i-med.ac.at>:
Hello cephers,
we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again.
We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby)
It started this night, I got the first HEALTH_WARN emails saying:
HEALTH_WARN
--- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
then it went on with:
HEALTH_WARN
--- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
--- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
Then one after another MDS was going to error state:
HEALTH_WARN
--- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state
=== Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more
In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1
ceph fs set cephfs max_mds 1
And set the filesystem down
ceph fs set cephfs down true
I tried to restart the MDS again but now I'm stuck at the following status:
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M)
data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep
[root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The end log file of the replay(laggy) default.cephmon-02.duujba shows:
[...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h:
In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h:
568: FAILED ceph_assert(p->first <= start)
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone()
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events ---
I have no idea how to resolve this and would be grateful for any help.
Dietmar _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Joachim, I suppose that setting the filesystem down will block all clients: ceph fs set cephfs down true right? Dietmar On 6/19/24 10:02, Joachim Kraftmayer wrote:
Hi Dietmar,
have you already blocked all cephfs clients?
Joachim
*Joachim Kraftmayer* CEO | p: +49 89 2152527-21 | e: joachim.kraftmayer@clyso.com <mailto:joachim.kraftmayer@clyso.com>
a: Loristr. 8 | 80335 Munich | Germany | w: https://clyso.com <https://clyso.com> | Utting a. A. | HR: Augsburg | HRB 25866 | USt. ID: DE275430677
Am Mi., 19. Juni 2024 um 09:44 Uhr schrieb Dietmar Rieder <dietmar.rieder@i-med.ac.at <mailto:dietmar.rieder@i-med.ac.at>>:
Hello cephers,
we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again.
We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby)
It started this night, I got the first HEALTH_WARN emails saying:
HEALTH_WARN
--- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
then it went on with:
HEALTH_WARN
--- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
--- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
Then one after another MDS was going to error state:
HEALTH_WARN
--- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state
=== Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more
In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1
ceph fs set cephfs max_mds 1
And set the filesystem down
ceph fs set cephfs down true
I tried to restart the MDS again but now I'm stuck at the following status:
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M)
data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep
[root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The end log file of the replay(laggy) default.cephmon-02.duujba shows:
[...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: 568: FAILED ceph_assert(p->first <= start)
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone()
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events ---
I have no idea how to resolve this and would be grateful for any help.
Dietmar _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
.... I did that after I have seen that it seems to be more severe, see my action "log" below. Dietmar On 6/19/24 10:05, Dietmar Rieder wrote:
Hi Joachim,
I suppose that setting the filesystem down will block all clients:
ceph fs set cephfs down true
right?
Dietmar
On 6/19/24 10:02, Joachim Kraftmayer wrote:
Hi Dietmar,
have you already blocked all cephfs clients?
Joachim
*Joachim Kraftmayer* CEO | p: +49 89 2152527-21 | e: joachim.kraftmayer@clyso.com <mailto:joachim.kraftmayer@clyso.com>
a: Loristr. 8 | 80335 Munich | Germany | w: https://clyso.com <https://clyso.com> | Utting a. A. | HR: Augsburg | HRB 25866 | USt. ID: DE275430677
Am Mi., 19. Juni 2024 um 09:44 Uhr schrieb Dietmar Rieder <dietmar.rieder@i-med.ac.at <mailto:dietmar.rieder@i-med.ac.at>>:
Hello cephers,
we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again.
We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby)
It started this night, I got the first HEALTH_WARN emails saying:
HEALTH_WARN
--- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
then it went on with:
HEALTH_WARN
--- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
--- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
Then one after another MDS was going to error state:
HEALTH_WARN
--- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state
=== Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more
In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1
ceph fs set cephfs max_mds 1
And set the filesystem down
ceph fs set cephfs down true
I tried to restart the MDS again but now I'm stuck at the following status:
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M)
data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep
[root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The end log file of the replay(laggy) default.cephmon-02.duujba shows:
[...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1
/home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000
/home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: 568: FAILED ceph_assert(p->first <= start)
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone()
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events ---
I have no idea how to resolve this and would be grateful for any help.
Dietmar _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
Hi Dietmar, On Wed, Jun 19, 2024 at 3:44 AM Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
Hello cephers,
we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again.
We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby)
It started this night, I got the first HEALTH_WARN emails saying:
HEALTH_WARN
--- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
then it went on with:
HEALTH_WARN
--- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
--- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
Then one after another MDS was going to error state:
HEALTH_WARN
--- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state
=== Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more
In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1
ceph fs set cephfs max_mds 1
This will not have any positive effect.
And set the filesystem down
ceph fs set cephfs down true
I tried to restart the MDS again but now I'm stuck at the following status:
Setting the file system "down" wont' do anything here either. What were you trying to accomplish? Restarting the MDS may only add to your problems.
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M)
data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep
[root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The end log file of the replay(laggy) default.cephmon-02.duujba shows:
[...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: 568: FAILED ceph_assert(p->first <= start)
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone()
Suggest following the recommendations by Xiubo. -- Patrick Donnelly, Ph.D. He / Him / His Red Hat Partner Engineer IBM, Inc. GPG: 19F28A586F808C2402351B93C3301A3E258DD79D
Hi Patrick, thanks for your message, see my comments below. (BTW it seem that there is an issue with the ceph mailing list, my previous message did not go through yet, so this may be redundant) On 6/19/24 17:27, Patrick Donnelly wrote:
Hi Dietmar,
On Wed, Jun 19, 2024 at 3:44 AM Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
Hello cephers,
we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again.
We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby)
It started this night, I got the first HEALTH_WARN emails saying:
HEALTH_WARN
--- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
then it went on with:
HEALTH_WARN
--- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
--- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
Then one after another MDS was going to error state:
HEALTH_WARN
--- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state
=== Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more
In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1
ceph fs set cephfs max_mds 1
This will not have any positive effect.
And set the filesystem down
ceph fs set cephfs down true
I tried to restart the MDS again but now I'm stuck at the following status:
Setting the file system "down" wont' do anything here either. What were you trying to accomplish? Restarting the MDS may only add to your problems.
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M)
data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep
[root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The end log file of the replay(laggy) default.cephmon-02.duujba shows:
[...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: 568: FAILED ceph_assert(p->first <= start)
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone()
Suggest following the recommendations by Xiubo.
I ran the disaster recovery procedures now as suggested by Xiubo, as follows: first I exported all the the journals, then I did [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 event recover_dentries summary Events by type: OPEN: 8737 PURGED: 1 SESSION: 9 SESSIONS: 2 SUBTREEMAP: 128 TABLECLIENT: 2 TABLESERVER: 30 UPDATE: 9207 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 event recover_dentries summary Events by type: OPEN: 3 SESSION: 1 SUBTREEMAP: 34 UPDATE: 32965 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 event recover_dentries summary Events by type: OPEN: 5289 SESSION: 10 SESSIONS: 3 SUBTREEMAP: 128 UPDATE: 76448 Errors: 0 [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:all journal inspect Overall journal integrity: OK Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 journal inspect Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 journal inspect Overall journal integrity: OK [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:0 journal reset old journal was 879331755046~508520587 new journal start will be 879843344384 (3068751 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal reset old journal was 934711229813~120432327 new journal start will be 934834864128 (3201988 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:2 journal reset old journal was 1334153584288~252692691 new journal start will be 1334409428992 (3152013 bytes past old end) writing journal head writing EResetJournal entry done [root@ceph01-b /]# cephfs-table-tool all reset session { "0": { "data": {}, "result": 0 }, "1": { "data": {}, "result": 0 }, "2": { "data": {}, "result": 0 } } [root@ceph01-b /]# cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: OK [root@ceph01-b /]# ceph fs reset cephfs --yes-i-really-mean-it But now I hit the error below: -20> 2024-06-19T11:13:00.610+0000 7ff3694d0700 10 monclient: _send_mon_message to mon.cephmon-03 at v2:10.1.3.23:3300/0 -19> 2024-06-19T11:13:00.637+0000 7ff3664ca700 2 mds.0.cache Memory usage: total 485928, rss 170860, heap 207156, baseline 182580, 0 / 33434 inodes have caps, 0 caps, 0 caps per inode -18> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.default.cephmon-03.chjusj Updating MDS map to version 8061 from mon.1 -17> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 handle_mds_map i am now mds.0.8058 -16> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 handle_mds_map state change up:rejoin --> up:active -15> 2024-06-19T11:13:00.787+0000 7ff36a4d2700 1 mds.0.8058 recovery_done -- successful recovery! -14> 2024-06-19T11:13:00.788+0000 7ff36a4d2700 1 mds.0.8058 active_start -13> 2024-06-19T11:13:00.789+0000 7ff36dcd9700 5 mds.beacon.default.cephmon-03.chjusj received beacon reply up:active seq 4 rtt 0.955007 -12> 2024-06-19T11:13:00.790+0000 7ff36a4d2700 1 mds.0.8058 cluster recovered. -11> 2024-06-19T11:13:00.790+0000 7ff36a4d2700 4 mds.0.8058 set_osd_epoch_barrier: epoch=33596 -10> 2024-06-19T11:13:00.790+0000 7ff3634c4700 5 mds.0.log _submit_thread 879843344432~2609 : EUpdate check_inode_max_size [metablob 0x100, 2 dirs] -9> 2024-06-19T11:13:00.791+0000 7ff3644c6700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/mds/MDCache.cc: In function 'void MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)' thread 7ff3644c6700 time 2024-06-19T11:13:00.791580+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/mds/MDCache.cc: 1660: FAILED ceph_assert(follows >= realm->get_newest_seq()) ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7ff374ad3e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7ff374ad3fdb] 3: (MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)+0x13c7) [0x55da0a7aa227] 4: (MDCache::journal_dirty_inode(MutationImpl*, EMetaBlob*, CInode*, snapid_t)+0xc5) [0x55da0a7aa3a5] 5: (Locker::check_inode_max_size(CInode*, bool, unsigned long, unsigned long, utime_t)+0x84d) [0x55da0a88ce3d] 6: (RecoveryQueue::_recovered(CInode*, int, unsigned long, utime_t)+0x4f0) [0x55da0a85ad50] 7: (MDSContext::complete(int)+0x5f) [0x55da0a9ddeef] 8: (MDSIOContextBase::complete(int)+0x524) [0x55da0a9de674] 9: (Filer::C_Probe::finish(int)+0xbb) [0x55da0aa9dc9b] 10: (Context::complete(int)+0xd) [0x55da0a6775fd] 11: (Finisher::finisher_thread_entry()+0x18d) [0x7ff374b77abd] 12: /lib64/libpthread.so.0(+0x81ca) [0x7ff3738791ca] 13: clone() -8> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client handle_log_ack log(last 7) v1 -7> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.647346+0000 mds.default.cephmon-03.chjusj (mds.0) 1 : cluster [ERR] loaded dup inode 0x10003e45d99 [415,head] v61632 at /home/balaz/.bash_history-54696.tmp, but inode 0x10003e45d99.head v61639 already exists at /home/balaz/.bash_history -6> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.648139+0000 mds.default.cephmon-03.chjusj (mds.0) 2 : cluster [ERR] loaded dup inode 0x10003e45d7c [415,head] v253612 at /home/rieder/.bash_history-10215.tmp, but inode 0x10003e45d7c.head v253630 already exists at /home/rieder/.bash_history -5> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.649483+0000 mds.default.cephmon-03.chjusj (mds.0) 3 : cluster [ERR] loaded dup inode 0x10003e45d83 [415,head] v164103 at /home/gottschling/.bash_history-44802.tmp, but inode 0x10003e45d83.head v164112 already exists at /home/gottschling/.bash_history -4> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.656221+0000 mds.default.cephmon-03.chjusj (mds.0) 4 : cluster [ERR] bad backtrace on directory inode 0x10003e42340 -3> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.737282+0000 mds.default.cephmon-03.chjusj (mds.0) 5 : cluster [ERR] bad backtrace on directory inode 0x10003e45d8b -2> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.804984+0000 mds.default.cephmon-03.chjusj (mds.0) 6 : cluster [ERR] bad backtrace on directory inode 0x10003e45d9f -1> 2024-06-19T11:13:00.792+0000 7ff36a4d2700 10 log_client logged 2024-06-19T11:12:59.805078+0000 mds.default.cephmon-03.chjusj (mds.0) 7 : cluster [ERR] bad backtrace on directory inode 0x10003e45d90 0> 2024-06-19T11:13:00.792+0000 7ff3644c6700 -1 *** Caught signal (Aborted) ** in thread 7ff3644c6700 thread_name:MR_Finisher ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7ff373883d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7ff374ad3e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7ff374ad3fdb] 6: (MDCache::journal_cow_dentry(MutationImpl*, EMetaBlob*, CDentry*, snapid_t, CInode**, CDentry::linkage_t*)+0x13c7) [0x55da0a7aa227] 7: (MDCache::journal_dirty_inode(MutationImpl*, EMetaBlob*, CInode*, snapid_t)+0xc5) [0x55da0a7aa3a5] 8: (Locker::check_inode_max_size(CInode*, bool, unsigned long, unsigned long, utime_t)+0x84d) [0x55da0a88ce3d] 9: (RecoveryQueue::_recovered(CInode*, int, unsigned long, utime_t)+0x4f0) [0x55da0a85ad50] 10: (MDSContext::complete(int)+0x5f) [0x55da0a9ddeef] 11: (MDSIOContextBase::complete(int)+0x524) [0x55da0a9de674] 12: (Filer::C_Probe::finish(int)+0xbb) [0x55da0aa9dc9b] 13: (Context::complete(int)+0xd) [0x55da0a6775fd] 14: (Finisher::finisher_thread_entry()+0x18d) [0x7ff374b77abd] 15: /lib64/libpthread.so.0(+0x81ca) [0x7ff3738791ca] 16: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this. --- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7ff362cc3700 / 7ff3634c4700 / md_submit 7ff363cc5700 / 7ff3644c6700 / MR_Finisher 7ff3654c8700 / PQ_Finisher 7ff365cc9700 / mds_rank_progr 7ff3664ca700 / ms_dispatch 7ff3684ce700 / ceph-mds 7ff3694d0700 / safe_timer 7ff36a4d2700 / ms_dispatch 7ff36b4d4700 / io_context_pool 7ff36c4d6700 / admin_socket 7ff36ccd7700 / msgr-worker-2 7ff36d4d8700 / msgr-worker-1 7ff36dcd9700 / msgr-worker-0 7ff375c9bb00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-03.chjusj.log --- end dump of recent events --- Any idea? Thanks Dietmar
Hi all, finally we were able to repair the filesystem and it seems that we did not lose any data. Thanks for all suggestions and comments. Here is a short summary of our journey: 1. At some point all our 6 MDS were going to error state one after another 2. We tried to restart them but they kept crashing 3. We learned that unfortunately we hit a known bug: <https://tracker.ceph.com/issues/61009> 4. We set the filesystem down "ceph fs set cephfs down true" and unmounted it from all clients. 5. We started with the disaster recovery procedure: <https://docs.ceph.com/en/reef/cephfs/disaster-recovery-experts/> I. Backup the journal cephfs-journal-tool --rank=cephfs:all journal export /mnt/backup/backup.bin II. DENTRY recovery from journal (We have 3 active MDS) cephfs-journal-tool --rank=cephfs:0 event recover_dentries summary cephfs-journal-tool --rank=cephfs:1 event recover_dentries summary cephfs-journal-tool --rank=cephfs:2 event recover_dentries summary cephfs-journal-tool --rank=cephfs:all journal inspect Overall journal integrity: OK Overall journal integrity: DAMAGED Corrupt regions: 0xd9a84f243c-ffffffffffffffff Overall journal integrity: OK The journal from rank 1 still shows damage III. Journal truncation cephfs-journal-tool --rank=cephfs:0 journal reset cephfs-journal-tool --rank=cephfs:1 journal reset cephfs-journal-tool --rank=cephfs:2 journal reset IV. MDS table wipes cephfs-table-tool all reset session cephfs-journal-tool --rank=cephfs:1 journal inspect Overall journal integrity: OK V. MDS MAP reset ceph fs reset cephfs --yes-i-really-mean-it After these steps to reset and trim the journal we tried to restart the MDS, however they were still dying shortly after starting. So as Xiubo suggested we went on with the disaster recovery procedure... VI. Recovery from missing metadata objects cephfs-table-tool 0 reset session cephfs-table-tool 0 reset snap cephfs-table-tool 0 reset inode cephfs-journal-tool --rank=cephfs:0 journal reset cephfs-data-scan init The "cephfs-data-scan init" gave us warnings about already existing inodes: Inode 0x0x1 already exists, skipping create. Use --force-init to overwrite the existing object. Inode 0x0x100 already exists, skipping create. Use --force-init to overwrite the existing object. We decided not to use --force-init and went on with cephfs-data-scan scan_extents sdd-rep-data-pool hdd-ec-data-pool The docs say it can take a "very long time", unfortunately the tool is not producing any ETA. After ~24hrs we interupted the process and restarted it with 32 workers. The parallel scan_extents took about 2h and 15 min and did not generate any output on stdout or stderr So we went on with parallel (32 workers) scan_inodes which also completed without any output after ~ 50 min. We then ran "cephfs-data-scan scan_links", however the tool was stopping afer ~ 45 min. with an error message: Error ((2) No such file or directory) We tried to go on anyway with "cephfs-data-scan cleanup". The cleanup was running for about 9h and 20min and did not produce any output. So we tried to startup the MDS again, however the still kept crashing: 2024-06-23T08:21:50.197+0000 7feeb5177700 1 mds.0.8075 rejoin_start 2024-06-23T08:21:50.201+0000 7feeb5177700 1 mds.0.8075 rejoin_joint_start 2024-06-23T08:21:50.204+0000 7feeaf16b700 1 mds.0.cache.den(0x10000000000 groups) loaded already corrupt dentry: [dentry #0x1/data/groups [bf,head] rep@0.0 NULL (dversion lock) pv=0 v=7910497 ino=(nil) state=0 0x55aa27a19180] [....] 2024-06-23T08:21:50.228+0000 7feeaf16b700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e42340 [...] 2024-06-23T08:21:50.345+0000 7feeaf16b700 -1 log_channel(cluster) log [ERR] : bad backtrace on directory inode 0x10003e45d8b [....] -6> 2024-06-23T08:21:50.351+0000 7feeaf16b700 10 log_client will send 2024-06-23T08:21:50.229835+0000 mds.default.cephmon-03.xcujhz (mds.0) 1 : cluster [ERR] bad backtrace on direc tory inode 0x10003e42340 -5> 2024-06-23T08:21:50.351+0000 7feeaf16b700 10 log_client will send 2024-06-23T08:21:50.347085+0000 mds.default.cephmon-03.xcujhz (mds.0) 2 : cluster [ERR] bad backtrace on directory inode 0x10003e45d8b -4> 2024-06-23T08:21:50.351+0000 7feeaf16b700 10 monclient: _send_mon_message to mon.cephmon-03 at v2:10.1.3.23:3300/0 -3> 2024-06-23T08:21:50.351+0000 7feeaf16b700 5 mds.beacon.default.cephmon-03.xcujhz Sending beacon down:damaged seq 90 -2> 2024-06-23T08:21:50.351+0000 7feeaf16b700 10 monclient: _send_mon_message to mon.cephmon-03 at v2:10.1.3.23:3300/0 -1> 2024-06-23T08:21:50.371+0000 7feeb817d700 5 mds.beacon.default.cephmon-03.xcujhz received beacon reply down:damaged seq 90 rtt 0.0200002 0> 2024-06-23T08:21:50.371+0000 7feeaf16b700 1 mds.default.cephmon-03.xcujhz respawn! So we decided to retry the "scan_links" and "cleanup" steps: cephfs-data-scan scan_links Took about 50min., no error this time. cephfs-data-scan cleanup Took about 10h, no error Now we again tried to fire up the MDS: we set "ceph mds repaired 0" and started the MDS. And now the cluster status was: HEALTH_OK However In the MDS logs we have seen some error messages: "bad backtrace on directory inode" VII. filesystem scrub We decided to run a filesystem scrub on one of the directories that showed these errors, which we identified by: rados --cluster ceph -p ssd-rep-metadata-pool listomapvals 10003e45d9f.00000000 ceph tell mds.cephfs:0 scrub start /directory/with/bad/backtrace recusive After this the cluser jumped to HEALTH_ERR state saying that there is a journal damage. We then decided to run a full filesystem scrub: ceph tell mds.cephfs:0 scrub start / recursive,repair,force It took about 4h to complete and from the logs we found about 192 "bad backtrace inodes" that did not have a corresponding "Scrub repaired inode" log message. We identified 2 home directories that were affected by this. We reran the the scrub recursive,repair,force on these two directories and ceph tell mds.cephfs:0 damage ls showed that the "bad backtrace inodes" were now reduced to 68 ("damage_type": "backtrace") We then ran "damage rm" on all these 68 IDs. ceph tell mds.cephfs:0 damage rm <ID> [...] After this, the cluster status went back to HEALTH_OK. VIII. final checkup Before we set the fileystem online again we ran a final checkup: ceph tell mds.cephfs:0 scrub start / recursive,repair,force The MDS log showed no more errors, so the filesystem and journal was consistent again. IX. new mount option for the clients and MDS settings to mitigate the bug In order not to be hit again by the bug (#61009) we set the -o wsync option for our kernel clients and mds_client_delegate_inos_pct 0 on our MDS. Now all is running fine again and we were lucky that no data was lost. X. Conclusion: If we would have be aware of the bug and its mitigation we would have saved a lot of downtime and some nerves. Is there an obvious place that I missed were such known issues are prominently made public? (The bug tracker maybe, but I think it is easy to miss the important among all others) Thanks again for all the help Dietmar On 6/19/24 09:43, Dietmar Rieder wrote:
Hello cephers,
we have a degraded filesystem on our ceph 18.2.2 cluster and I'd need to get it up again.
We have 6 MDS daemons and (3 active, each pinned to a subtree, 3 standby)
It started this night, I got the first HEALTH_WARN emails saying:
HEALTH_WARN
--- New --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
then it went on with:
HEALTH_WARN
--- New --- [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
--- Cleared --- [WARN] MDS_CLIENT_RECALL: 1 clients failing to respond to cache pressure mds.default.cephmon-02.duujba(mds.1): Client apollo-10:cephfs_user failing to respond to cache pressure client_id: 1962074
=== Full health status === [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded
Then one after another MDS was going to error state:
HEALTH_WARN
--- Updated --- [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state
=== Full health status === [WARN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in error state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WARN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WARN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more
In the morning then I tried to restart the MDS in error state but the kept failing. I then reduced the number of active MDS to 1
ceph fs set cephfs max_mds 1
And set the filesystem down
ceph fs set cephfs down true
I tried to restart the MDS again but now I'm stuck at the following status:
[root@ceph01-b ~]# ceph -s cluster: id: aae23c5c-a98b-11ee-b44d-00620b05cac4 health: HEALTH_WARN 4 failed cephadm daemon(s) 1 filesystem is degraded insufficient standby MDS daemons available
services: mon: 3 daemons, quorum cephmon-01,cephmon-03,cephmon-02 (age 2w) mgr: cephmon-01.dsxcho(active, since 11w), standbys: cephmon-02.nssigg, cephmon-03.rgefle mds: 3/3 daemons up osd: 336 osds: 336 up (since 11w), 336 in (since 3M)
data: volumes: 0/1 healthy, 1 recovering pools: 4 pools, 6401 pgs objects: 284.69M objects, 623 TiB usage: 889 TiB used, 3.1 PiB / 3.9 PiB avail pgs: 6186 active+clean 156 active+clean+scrubbing 59 active+clean+scrubbing+deep
[root@ceph01-b ~]# ceph health detail HEALTH_WARN 4 failed cephadm daemon(s); 1 filesystem is degraded; insufficient standby MDS daemons available [WRN] CEPHADM_FAILED_DAEMON: 4 failed cephadm daemon(s) daemon mds.default.cephmon-01.cepqjp on cephmon-01 is in error state daemon mds.default.cephmon-02.duujba on cephmon-02 is in unknown state daemon mds.default.cephmon-03.chjusj on cephmon-03 is in error state daemon mds.default.cephmon-03.xcujhz on cephmon-03 is in error state [WRN] FS_DEGRADED: 1 filesystem is degraded fs cephfs is degraded [WRN] MDS_INSUFFICIENT_STANDBY: insufficient standby MDS daemons available have 0; want 1 more [root@ceph01-b ~]# [root@ceph01-b ~]# ceph fs status cephfs - 40 clients ====== RANK STATE MDS ACTIVITY DNS INOS DIRS CAPS 0 resolve default.cephmon-02.nyfook 12.3k 11.8k 3228 0 1 replay(laggy) default.cephmon-02.duujba 0 0 0 0 2 resolve default.cephmon-01.pvnqad 15.8k 3541 1409 0 POOL TYPE USED AVAIL ssd-rep-metadata-pool metadata 295G 63.5T sdd-rep-data-pool data 10.2T 84.6T hdd-ec-data-pool data 808T 1929T MDS version: ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable)
The end log file of the replay(laggy) default.cephmon-02.duujba shows:
[...] -11> 2024-06-19T07:12:38.980+0000 7f90fd117700 1 mds.1.journaler.pq(ro) _finish_probe_end write_pos = 8673820672 (header had 8623488918). recovered. -10> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): open complete -9> 2024-06-19T07:12:38.980+0000 7f90fd117700 4 mds.1.purge_queue operator(): recovering write_pos -8> 2024-06-19T07:12:39.015+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ef42c00 auth_method 0 -7> 2024-06-19T07:12:39.025+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93ef43400 auth_method 0 -6> 2024-06-19T07:12:39.038+0000 7f90fd117700 4 mds.1.purge_queue _recover: write_pos recovered -5> 2024-06-19T07:12:39.038+0000 7f90fd117700 1 mds.1.journaler.pq(ro) set_writeable -4> 2024-06-19T07:12:39.044+0000 7f9105127700 10 monclient: get_auth_request con 0x55a93ef43c00 auth_method 0 -3> 2024-06-19T07:12:39.113+0000 7f9104926700 10 monclient: get_auth_request con 0x55a93ed97000 auth_method 0 -2> 2024-06-19T07:12:39.123+0000 7f9105928700 10 monclient: get_auth_request con 0x55a93e903c00 auth_method 0 -1> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: In function 'void interval_set<T, C>::erase(T, T, std::function<bool(T, T)>) [with T = inodeno_t; C = std::map]' thread 7f90fa912700 time 2024-06-19T07:12:39.235633+0000 /home/jenkins-build/build/workspace/ceph-build/ARCH/x86_64/AVAILABLE_ARCH/x86_64/AVAILABLE_DIST/centos8/DIST/centos8/MACHINE_SIZE/gigantic/release/18.2.2/rpm/el8/BUILD/ceph-18.2.2/src/include/interval_set.h: 568: FAILED ceph_assert(p->first <= start)
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x135) [0x7f910c722e15] 2: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 3: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 4: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 5: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 6: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 7: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 8: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 9: clone()
0> 2024-06-19T07:12:39.236+0000 7f90fa912700 -1 *** Caught signal (Aborted) ** in thread 7f90fa912700 thread_name:md_log_replay
ceph version 18.2.2 (531c0d11a1c5d39fbfe6aa8a521f023abf3bf3e2) reef (stable) 1: /lib64/libpthread.so.0(+0x12d20) [0x7f910b4d2d20] 2: gsignal() 3: abort() 4: (ceph::__ceph_assert_fail(char const*, char const*, int, char const*)+0x18f) [0x7f910c722e6f] 5: /usr/lib64/ceph/libceph-common.so.2(+0x2a9fdb) [0x7f910c722fdb] 6: (interval_set<inodeno_t, std::map>::erase(inodeno_t, inodeno_t, std::function<bool (inodeno_t, inodeno_t)>)+0x2e5) [0x55a93c0de9a5] 7: (EMetaBlob::replay(MDSRank*, LogSegment*, int, MDPeerUpdate*)+0x4207) [0x55a93c3e76e7] 8: (EUpdate::replay(MDSRank*)+0x61) [0x55a93c3e9f81] 9: (MDLog::_replay_thread()+0x6c9) [0x55a93c3701d9] 10: (MDLog::ReplayThread::entry()+0x11) [0x55a93c01e2d1] 11: /lib64/libpthread.so.0(+0x81ca) [0x7f910b4c81ca] 12: clone() NOTE: a copy of the executable, or `objdump -rdS <executable>` is needed to interpret this.
--- logging levels --- 0/ 5 none 0/ 1 lockdep 0/ 1 context 1/ 1 crush 1/ 5 mds 1/ 5 mds_balancer 1/ 5 mds_locker 1/ 5 mds_log 1/ 5 mds_log_expire 1/ 5 mds_migrator 0/ 1 buffer 0/ 1 timer 0/ 1 filer 0/ 1 striper 0/ 1 objecter 0/ 5 rados 0/ 5 rbd 0/ 5 rbd_mirror 0/ 5 rbd_replay 0/ 5 rbd_pwl 0/ 5 journaler 0/ 5 objectcacher 0/ 5 immutable_obj_cache 0/ 5 client 1/ 5 osd 0/ 5 optracker 0/ 5 objclass 1/ 3 filestore 1/ 3 journal 0/ 0 ms 1/ 5 mon 0/10 monc 1/ 5 paxos 0/ 5 tp 1/ 5 auth 1/ 5 crypto 1/ 1 finisher 1/ 1 reserver 1/ 5 heartbeatmap 1/ 5 perfcounter 1/ 5 rgw 1/ 5 rgw_sync 1/ 5 rgw_datacache 1/ 5 rgw_access 1/ 5 rgw_dbstore 1/ 5 rgw_flight 1/ 5 javaclient 1/ 5 asok 1/ 1 throttle 0/ 0 refs 1/ 5 compressor 1/ 5 bluestore 1/ 5 bluefs 1/ 3 bdev 1/ 5 kstore 4/ 5 rocksdb 4/ 5 leveldb 1/ 5 fuse 2/ 5 mgr 1/ 5 mgrc 1/ 5 dpdk 1/ 5 eventtrace 1/ 5 prioritycache 0/ 5 test 0/ 5 cephfs_mirror 0/ 5 cephsqlite 0/ 5 seastore 0/ 5 seastore_onode 0/ 5 seastore_odata 0/ 5 seastore_omap 0/ 5 seastore_tm 0/ 5 seastore_t 0/ 5 seastore_cleaner 0/ 5 seastore_epm 0/ 5 seastore_lba 0/ 5 seastore_fixedkv_tree 0/ 5 seastore_cache 0/ 5 seastore_journal 0/ 5 seastore_device 0/ 5 seastore_backref 0/ 5 alienstore 1/ 5 mclock 0/ 5 cyanstore 1/ 5 ceph_exporter 1/ 5 memstore -2/-2 (syslog threshold) -1/-1 (stderr threshold) --- pthread ID / name mapping for recent threads --- 7f90fa912700 / md_log_replay 7f90fb914700 / 7f90fc115700 / MR_Finisher 7f90fd117700 / PQ_Finisher 7f90fe119700 / ms_dispatch 7f910011d700 / ceph-mds 7f9102121700 / ms_dispatch 7f9103123700 / io_context_pool 7f9104125700 / admin_socket 7f9104926700 / msgr-worker-2 7f9105127700 / msgr-worker-1 7f9105928700 / msgr-worker-0 7f910d8eab00 / ceph-mds max_recent 10000 max_new 1000 log_file /var/log/ceph/ceph-mds.default.cephmon-02.duujba.log --- end dump of recent events ---
I have no idea how to resolve this and would be grateful for any help.
Dietmar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Dietmar, On 29-06-2024 10:50, Dietmar Rieder wrote:
Hi all,
finally we were able to repair the filesystem and it seems that we did not lose any data. Thanks for all suggestions and comments.
Here is a short summary of our journey:
Thanks for writing this up. This might be useful for someone in the future. --- snip ---
X. Conclusion:
If we would have be aware of the bug and its mitigation we would have saved a lot of downtime and some nerves.
Is there an obvious place that I missed were such known issues are prominently made public? (The bug tracker maybe, but I think it is easy to miss the important among all others)
Not that I know of. But changes in behavior of Ceph (daemons) and or Ceph kernels would be good to know about indeed. I follow the ceph-kernel mailing list to see what is going on with the development of kernel CephFS. And there is a thread about reverting the PR that Enrico linked to [1], here the last mail in that thread from Venky to Ilya [2]: "Hi Ilya, After some digging and talking to Jeff, I figured that it's possible to disable async dirops from the mds side by setting `mds_client_delegate_inos_pct` config to 0: - name: mds_client_delegate_inos_pct type: uint level: advanced desc: percentage of preallocated inos to delegate to client default: 50 services: - mds So, I guess this patch is really not required. We can suggest this config update to users and document it for now. We lack tests with this config disabled, so I'll be adding the same before recommending it out. Will keep you posted." However, I have not seen any update after this. So apparently it is possible to disable this preallocate behavior globally by disabling it on the MDS. But there are (were) no MDS tests with this option disabled (I guess a percentage of "0" would disable it). So I'm not sure it's safe to disable it, and what would happen if you disable this on the MDS when there are clients actually using preallocated inodes. I have added Venky in the CC so I hope he can give us an update about the recommended way(s) of disabling preallocated inodes Gr. Stefan [1]: https://github.com/gregkh/linux/commit/f7a67b463fb83a4b9b11ceaa8ec4950b8fb7f... [2]: https://lore.kernel.org/all/20231003110556.140317-1-vshankar@redhat.com/T/
Hi Stefan, On 7/1/24 10:34, Stefan Kooman wrote:
Hi Dietmar,
On 29-06-2024 10:50, Dietmar Rieder wrote:
Hi all,
finally we were able to repair the filesystem and it seems that we did not lose any data. Thanks for all suggestions and comments.
Here is a short summary of our journey:
Thanks for writing this up. This might be useful for someone in the future.
Yeah, your welcome, I thought so too
--- snip ---
X. Conclusion:
If we would have be aware of the bug and its mitigation we would have saved a lot of downtime and some nerves.
Is there an obvious place that I missed were such known issues are prominently made public? (The bug tracker maybe, but I think it is easy to miss the important among all others)
Not that I know of. But changes in behavior of Ceph (daemons) and or Ceph kernels would be good to know about indeed. I follow the ceph-kernel mailing list to see what is going on with the development of kernel CephFS. And there is a thread about reverting the PR that Enrico linked to [1], here the last mail in that thread from Venky to Ilya [2]:
"Hi Ilya,
After some digging and talking to Jeff, I figured that it's possible to disable async dirops from the mds side by setting `mds_client_delegate_inos_pct` config to 0:
- name: mds_client_delegate_inos_pct type: uint level: advanced desc: percentage of preallocated inos to delegate to client default: 50 services: - mds
So, I guess this patch is really not required. We can suggest this config update to users and document it for now. We lack tests with this config disabled, so I'll be adding the same before recommending it out. Will keep you posted."
However, I have not seen any update after this. So apparently it is possible to disable this preallocate behavior globally by disabling it on the MDS. But there are (were) no MDS tests with this option disabled (I guess a percentage of "0" would disable it). So I'm not sure it's safe to disable it, and what would happen if you disable this on the MDS when there are clients actually using preallocated inodes. I have added Venky in the CC so I hope he can give us an update about the recommended way(s) of disabling preallocated inodes
Gr. Stefan
[1]: https://github.com/gregkh/linux/commit/f7a67b463fb83a4b9b11ceaa8ec4950b8fb7f...
[2]: https://lore.kernel.org/all/20231003110556.140317-1-vshankar@redhat.com/T/
I'm curious about any updates as well. I hope that not too many cephfs users will end up in this situation until the bug is fixed furthermore I hope that with posting our experiences here some will get alerted and make the proposed mitigation settings.... Best Dietmar
Hi Stefan, On Mon, Jul 1, 2024 at 2:30 PM Stefan Kooman <stefan@bit.nl> wrote:
Hi Dietmar,
On 29-06-2024 10:50, Dietmar Rieder wrote:
Hi all,
finally we were able to repair the filesystem and it seems that we did not lose any data. Thanks for all suggestions and comments.
Here is a short summary of our journey:
Thanks for writing this up. This might be useful for someone in the future.
--- snip ---
X. Conclusion:
If we would have be aware of the bug and its mitigation we would have saved a lot of downtime and some nerves.
Is there an obvious place that I missed were such known issues are prominently made public? (The bug tracker maybe, but I think it is easy to miss the important among all others)
Not that I know of. But changes in behavior of Ceph (daemons) and or Ceph kernels would be good to know about indeed. I follow the ceph-kernel mailing list to see what is going on with the development of kernel CephFS. And there is a thread about reverting the PR that Enrico linked to [1], here the last mail in that thread from Venky to Ilya [2]:
"Hi Ilya,
After some digging and talking to Jeff, I figured that it's possible to disable async dirops from the mds side by setting `mds_client_delegate_inos_pct` config to 0:
- name: mds_client_delegate_inos_pct type: uint level: advanced desc: percentage of preallocated inos to delegate to client default: 50 services: - mds
So, I guess this patch is really not required. We can suggest this config update to users and document it for now. We lack tests with this config disabled, so I'll be adding the same before recommending it out. Will keep you posted."
However, I have not seen any update after this. So apparently it is possible to disable this preallocate behavior globally by disabling it on the MDS. But there are (were) no MDS tests with this option disabled (I guess a percentage of "0" would disable it). So I'm not sure it's safe to disable it, and what would happen if you disable this on the MDS when there are clients actually using preallocated inodes. I have added Venky in the CC so I hope he can give us an update about the recommended way(s) of disabling preallocated inodes
It's safe to disable preallocation by setting `mds_client_delegate_inos_pct = 0'. Once disabled, the MDS will not delegate (preallocated) inode ranges to clients, effectively disabling async dirops. We have seen users running with this config (and using the `wsync' mount option in the kernel driver - although setting both isn't really required IMO, Xiubo?) and reporting stabe file system operations. As far as tests are concerned, the way forward is to have a shot reproducing this in our test lab. We have a tracker for this: https://tracker.ceph.com/issues/66250 It's likely that a combination of having the MDS preallocate inodes and delegate a percentage of those frequently to clients is causing this bug. Furthermore, an enhancement is proposed (for the shorter term) to not crash the MDS, but to blocklist the client that's holding the "problematic" preallocated inode range. That way, the file system isn't totally unavailable when such a problem occurs (the client would have to be remounted though, but that's a lesser pain than going through the disaster recovery steps). HTH.
Gr. Stefan
[1]: https://github.com/gregkh/linux/commit/f7a67b463fb83a4b9b11ceaa8ec4950b8fb7f...
[2]: https://lore.kernel.org/all/20231003110556.140317-1-vshankar@redhat.com/T/
-- Cheers, Venky
Hi Venky, On 02-07-2024 09:45, Venky Shankar wrote:
Hi Stefan,
On Mon, Jul 1, 2024 at 2:30 PM Stefan Kooman <stefan@bit.nl> wrote:
Hi Dietmar,
On 29-06-2024 10:50, Dietmar Rieder wrote:
Hi all,
finally we were able to repair the filesystem and it seems that we did not lose any data. Thanks for all suggestions and comments.
Here is a short summary of our journey:
Thanks for writing this up. This might be useful for someone in the future.
--- snip ---
X. Conclusion:
If we would have be aware of the bug and its mitigation we would have saved a lot of downtime and some nerves.
Is there an obvious place that I missed were such known issues are prominently made public? (The bug tracker maybe, but I think it is easy to miss the important among all others)
Not that I know of. But changes in behavior of Ceph (daemons) and or Ceph kernels would be good to know about indeed. I follow the ceph-kernel mailing list to see what is going on with the development of kernel CephFS. And there is a thread about reverting the PR that Enrico linked to [1], here the last mail in that thread from Venky to Ilya [2]:
"Hi Ilya,
After some digging and talking to Jeff, I figured that it's possible to disable async dirops from the mds side by setting `mds_client_delegate_inos_pct` config to 0:
- name: mds_client_delegate_inos_pct type: uint level: advanced desc: percentage of preallocated inos to delegate to client default: 50 services: - mds
So, I guess this patch is really not required. We can suggest this config update to users and document it for now. We lack tests with this config disabled, so I'll be adding the same before recommending it out. Will keep you posted."
However, I have not seen any update after this. So apparently it is possible to disable this preallocate behavior globally by disabling it on the MDS. But there are (were) no MDS tests with this option disabled (I guess a percentage of "0" would disable it). So I'm not sure it's safe to disable it, and what would happen if you disable this on the MDS when there are clients actually using preallocated inodes. I have added Venky in the CC so I hope he can give us an update about the recommended way(s) of disabling preallocated inodes
It's safe to disable preallocation by setting `mds_client_delegate_inos_pct = 0'. Once disabled, the MDS will not delegate (preallocated) inode ranges to clients, effectively disabling async dirops. We have seen users running with this config (and using the `wsync' mount option in the kernel driver - although setting both isn't really required IMO, Xiubo?) and reporting stabe file system operations.
Can this be done live, as in via `ceph config set mds mds_client_delegate_inos_pct = 0` and / or `ceph daemon mds.$hostname config set mds_client_delegate_inos_pct = 0`? Has that been tested? Or is it safer to do this by restarting the MDS? I wonder how the MDS handles cases where inodes are already delegated to the client and has to transition to full sync behavior again.
As far as tests are concerned, the way forward is to have a shot reproducing this in our test lab. We have a tracker for this:
https://tracker.ceph.com/issues/66250
It's likely that a combination of having the MDS preallocate inodes and delegate a percentage of those frequently to clients is causing this bug. Furthermore, an enhancement is proposed (for the shorter term) to not crash the MDS, but to blocklist the client that's holding the "problematic" preallocated inode range. That way, the file system isn't totally unavailable when such a problem occurs (the client would have to be remounted though, but that's a lesser pain than going through the disaster recovery steps).
Gr. Stefan
Hi Stefan, On Tue, Jul 2, 2024 at 4:16 PM Stefan Kooman <stefan@bit.nl> wrote:
Hi Venky,
On 02-07-2024 09:45, Venky Shankar wrote:
Hi Stefan,
On Mon, Jul 1, 2024 at 2:30 PM Stefan Kooman <stefan@bit.nl> wrote:
Hi Dietmar,
On 29-06-2024 10:50, Dietmar Rieder wrote:
Hi all,
finally we were able to repair the filesystem and it seems that we did not lose any data. Thanks for all suggestions and comments.
Here is a short summary of our journey:
Thanks for writing this up. This might be useful for someone in the future.
--- snip ---
X. Conclusion:
If we would have be aware of the bug and its mitigation we would have saved a lot of downtime and some nerves.
Is there an obvious place that I missed were such known issues are prominently made public? (The bug tracker maybe, but I think it is easy to miss the important among all others)
Not that I know of. But changes in behavior of Ceph (daemons) and or Ceph kernels would be good to know about indeed. I follow the ceph-kernel mailing list to see what is going on with the development of kernel CephFS. And there is a thread about reverting the PR that Enrico linked to [1], here the last mail in that thread from Venky to Ilya [2]:
"Hi Ilya,
After some digging and talking to Jeff, I figured that it's possible to disable async dirops from the mds side by setting `mds_client_delegate_inos_pct` config to 0:
- name: mds_client_delegate_inos_pct type: uint level: advanced desc: percentage of preallocated inos to delegate to client default: 50 services: - mds
So, I guess this patch is really not required. We can suggest this config update to users and document it for now. We lack tests with this config disabled, so I'll be adding the same before recommending it out. Will keep you posted."
However, I have not seen any update after this. So apparently it is possible to disable this preallocate behavior globally by disabling it on the MDS. But there are (were) no MDS tests with this option disabled (I guess a percentage of "0" would disable it). So I'm not sure it's safe to disable it, and what would happen if you disable this on the MDS when there are clients actually using preallocated inodes. I have added Venky in the CC so I hope he can give us an update about the recommended way(s) of disabling preallocated inodes
It's safe to disable preallocation by setting `mds_client_delegate_inos_pct = 0'. Once disabled, the MDS will not delegate (preallocated) inode ranges to clients, effectively disabling async dirops. We have seen users running with this config (and using the `wsync' mount option in the kernel driver - although setting both isn't really required IMO, Xiubo?) and reporting stabe file system operations.
Can this be done live, as in via `ceph config set mds mds_client_delegate_inos_pct = 0` and / or `ceph daemon mds.$hostname config set mds_client_delegate_inos_pct = 0`?
Yes - use `config set`.
Has that been tested?
Somewhat yes - it's safe to disable the config.
Or is it safer to do this by restarting the MDS? I wonder how the MDS handles cases where inodes are already delegated to the client and has to transition to full sync behavior again.
The MDS would use the updated config on subsequent client requests. The already preallocated inodes would be continued to be used by the client till exhaustion.
As far as tests are concerned, the way forward is to have a shot reproducing this in our test lab. We have a tracker for this:
https://tracker.ceph.com/issues/66250
It's likely that a combination of having the MDS preallocate inodes and delegate a percentage of those frequently to clients is causing this bug. Furthermore, an enhancement is proposed (for the shorter term) to not crash the MDS, but to blocklist the client that's holding the "problematic" preallocated inode range. That way, the file system isn't totally unavailable when such a problem occurs (the client would have to be remounted though, but that's a lesser pain than going through the disaster recovery steps).
Gr. Stefan
-- Cheers, Venky
Hi, On 01-07-2024 10:34, Stefan Kooman wrote:
Not that I know of. But changes in behavior of Ceph (daemons) and or Ceph kernels would be good to know about indeed. I follow the ceph-kernel mailing list to see what is going on with the development of kernel CephFS. And there is a thread about reverting the PR that Enrico linked to [1], here the last mail in that thread from Venky to Ilya [2]:
One more thing: The man page (mount.ceph) incorrectly states that "wsync" is the default [1], which it isn't anymore since 5.16 [2]. Not sure who is responsible for maintaining the man page, so CC'ing Zac and Ceph kernel developers. I would add information in the man page what options are default in what kernels (and to keep the man page up to date in the future). Gr. Stefan [1]: https://docs.ceph.com/en/latest/man/8/mount.ceph/ [2]: https://github.com/gregkh/linux/commit/f7a67b463fb83a4b9b11ceaa8ec4950b8fb7f...
participants (7)
-
Dietmar Rieder
-
Enrico Bocchi
-
Joachim Kraftmayer
-
Patrick Donnelly
-
Stefan Kooman
-
Venky Shankar
-
Xiubo Li