MDS stuck in up:stopping state
After scaling the number of MDS daemons down, we now have a daemon stuck in the "up:stopping" state. The documentation says it can take several minutes to stop the daemon, but it has been stuck in this state for almost a full day. According to the "ceph fs status" output attached below, it still holds information about 2 inodes, which we assume is the reason why it cannot stop completely. Does anyone know what we can do to finally stop it? cephfs - 71 clients ====== RANK STATE MDS ACTIVITY DNS INOS 0 active ceph-mon-01 Reqs: 0 /s 15.7M 15.4M 1 active ceph-mon-02 Reqs: 48 /s 19.7M 17.1M 2 stopping ceph-mon-03 0 2 POOL TYPE USED AVAIL cephfs_metadata metadata 652G 185T cephfs_data data 1637T 539T STANDBY MDS ceph-mon-03-mds-2 MDS version: ceph version 15.2.11 (e3523634d9c2227df9af89a4eac33d16738c49cb) octopus (stable)
Hi Martin, You may hit https://tracker.ceph.com/issues/50112, which we failed to find the root cause yet. I resolved this by restart rank 0. (I have only 2 active MDSs) Weiwen Hu 发送自 Windows 10 版邮件<https://go.microsoft.com/fwlink/?LinkId=550986>应用 发件人: Martin Rasmus Lundquist Hansen<mailto:hansen@imada.sdu.dk> 发送时间: 2021年5月27日 14:26 收件人: ceph-users@ceph.io<mailto:ceph-users@ceph.io> 主题: [ceph-users] MDS stuck in up:stopping state After scaling the number of MDS daemons down, we now have a daemon stuck in the "up:stopping" state. The documentation says it can take several minutes to stop the daemon, but it has been stuck in this state for almost a full day. According to the "ceph fs status" output attached below, it still holds information about 2 inodes, which we assume is the reason why it cannot stop completely. Does anyone know what we can do to finally stop it? cephfs - 71 clients ====== RANK STATE MDS ACTIVITY DNS INOS 0 active ceph-mon-01 Reqs: 0 /s 15.7M 15.4M 1 active ceph-mon-02 Reqs: 48 /s 19.7M 17.1M 2 stopping ceph-mon-03 0 2 POOL TYPE USED AVAIL cephfs_metadata metadata 652G 185T cephfs_data data 1637T 539T STANDBY MDS ceph-mon-03-mds-2 MDS version: ceph version 15.2.11 (e3523634d9c2227df9af89a4eac33d16738c49cb) octopus (stable) _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On Thu, May 27, 2021 at 07:02:16AM +0000, 胡 玮文 wrote:
You may hit https://tracker.ceph.com/issues/50112, which we failed to find the root cause yet. I resolved this by restart rank 0. (I have only 2 active MDSs)
I have this exact issue while trying to upgrade from 12.2 (which is pending this mds issue). I don't have any active clients, restarting rank0 does not help. +------+----------+-----------+---------------+-------+-------+ | Rank | State | MDS | Activity | dns | inos | +------+----------+-----------+---------------+-------+-------+ | 0 | active | osdnode05 | Reqs: 0 /s | 2760k | 2760k | | 1 | stopping | osdnode06 | | 10 | 11 | +------+----------+-----------+---------------+-------+-------+ -- Mark Schouten | Tuxis B.V. KvK: 74698818 | http://www.tuxis.nl/ T: +31 318 200208 | info@tuxis.nl
On Thu, May 27, 2021 at 10:37:33AM +0200, Mark Schouten wrote:
On Thu, May 27, 2021 at 07:02:16AM +0000, 胡 玮文 wrote:
You may hit https://tracker.ceph.com/issues/50112, which we failed to find the root cause yet. I resolved this by restart rank 0. (I have only 2 active MDSs)
I have this exact issue while trying to upgrade from 12.2 (which is pending this mds issue). I don't have any active clients, restarting rank0 does not help.
Since I have no active clients. Can I just shut down the all mds'es, upgrade them and expect an upgrade to fix this magically? Or would upgrading possibly break the CephFS-fs? -- Mark Schouten | Tuxis B.V. KvK: 74698818 | http://www.tuxis.nl/ T: +31 318 200208 | info@tuxis.nl
Hi Weiwen, Amazing, that actually worked. So simple, thanks! ________________________________ Fra: 胡 玮文 <huww98@outlook.com> Sendt: 27. maj 2021 09:02 Til: Martin Rasmus Lundquist Hansen <hansen@imada.sdu.dk>; ceph-users@ceph.io <ceph-users@ceph.io> Emne: 回复: MDS stuck in up:stopping state Hi Martin, You may hit https://tracker.ceph.com/issues/50112<https://eur03.safelinks.protection.outlook.com/?url=https%3A%2F%2Ftracker.ceph.com%2Fissues%2F50112&data=04%7C01%7Chansen%40imada.sdu.dk%7C557b70c448af40a9b9ad08d920dd5fa8%7C9a97c27db83e4694b35354bdbf18ab5b%7C0%7C0%7C637576958554129710%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C2000&sdata=ztjzyQDiHek1F6BA3mCcRHPKkGPuWheZ0Pi0NcK6Xmw%3D&reserved=0>, which we failed to find the root cause yet. I resolved this by restart rank 0. (I have only 2 active MDSs) Weiwen Hu 发送自 Windows 10 版邮件<https://eur03.safelinks.protection.outlook.com/?url=https%3A%2F%2Fgo.microsoft.com%2Ffwlink%2F%3FLinkId%3D550986&data=04%7C01%7Chansen%40imada.sdu.dk%7C557b70c448af40a9b9ad08d920dd5fa8%7C9a97c27db83e4694b35354bdbf18ab5b%7C0%7C0%7C637576958554139710%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C2000&sdata=2VVst65FTtm0O%2FX5h%2F%2BLeg51%2F5UbLNASibwKPnRM7%2Fo%3D&reserved=0>应用 发件人: Martin Rasmus Lundquist Hansen<mailto:hansen@imada.sdu.dk> 发送时间: 2021年5月27日 14:26 收件人: ceph-users@ceph.io<mailto:ceph-users@ceph.io> 主题: [ceph-users] MDS stuck in up:stopping state After scaling the number of MDS daemons down, we now have a daemon stuck in the "up:stopping" state. The documentation says it can take several minutes to stop the daemon, but it has been stuck in this state for almost a full day. According to the "ceph fs status" output attached below, it still holds information about 2 inodes, which we assume is the reason why it cannot stop completely. Does anyone know what we can do to finally stop it? cephfs - 71 clients ====== RANK STATE MDS ACTIVITY DNS INOS 0 active ceph-mon-01 Reqs: 0 /s 15.7M 15.4M 1 active ceph-mon-02 Reqs: 48 /s 19.7M 17.1M 2 stopping ceph-mon-03 0 2 POOL TYPE USED AVAIL cephfs_metadata metadata 652G 185T cephfs_data data 1637T 539T STANDBY MDS ceph-mon-03-mds-2 MDS version: ceph version 15.2.11 (e3523634d9c2227df9af89a4eac33d16738c49cb) octopus (stable) _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On Thu, May 27, 2021 at 06:25:44AM +0000, Martin Rasmus Lundquist Hansen wrote:
After scaling the number of MDS daemons down, we now have a daemon stuck in the "up:stopping" state. The documentation says it can take several minutes to stop the daemon, but it has been stuck in this state for almost a full day. According to the "ceph fs status" output attached below, it still holds information about 2 inodes, which we assume is the reason why it cannot stop completely.
Does anyone know what we can do to finally stop it?
I have no clients, and it still does not want to stop rank1. Funny thing is, while trying to fix this by restarting mdses, I sometimes see a list of clients popping up in the dashboard, even though no clients are connected.. -- Mark Schouten | Tuxis B.V. KvK: 74698818 | http://www.tuxis.nl/ T: +31 318 200208 | info@tuxis.nl
On Thu, May 27, 2021 at 12:38:07PM +0200, Mark Schouten wrote:
On Thu, May 27, 2021 at 06:25:44AM +0000, Martin Rasmus Lundquist Hansen wrote:
After scaling the number of MDS daemons down, we now have a daemon stuck in the "up:stopping" state. The documentation says it can take several minutes to stop the daemon, but it has been stuck in this state for almost a full day. According to the "ceph fs status" output attached below, it still holds information about 2 inodes, which we assume is the reason why it cannot stop completely.
Does anyone know what we can do to finally stop it?
I have no clients, and it still does not want to stop rank1. Funny thing is, while trying to fix this by restarting mdses, I sometimes see a list of clients popping up in the dashboard, even though no clients are connected..
Configuring debuglogging shows me the following: https://p.6core.net/p/rlMaunS8IM1AY5E58uUB6oy4 I have quite a lot of hardlinks on this filesystem, which I've seen issue with 'No space left on device'. I have mds_bal_fragment_size_max set to 200000 to mitigate that. The message 'waiting for strays to migrate' makes me feel like I should push the MDS to migrate them somehow .. But how? -- Mark Schouten | Tuxis B.V. KvK: 74698818 | http://www.tuxis.nl/ T: +31 318 200208 | info@tuxis.nl
在 2021年5月27日,19:11,Mark Schouten <mark@tuxis.nl> 写道:
On Thu, May 27, 2021 at 12:38:07PM +0200, Mark Schouten wrote:
On Thu, May 27, 2021 at 06:25:44AM +0000, Martin Rasmus Lundquist Hansen wrote: After scaling the number of MDS daemons down, we now have a daemon stuck in the "up:stopping" state. The documentation says it can take several minutes to stop the daemon, but it has been stuck in this state for almost a full day. According to the "ceph fs status" output attached below, it still holds information about 2 inodes, which we assume is the reason why it cannot stop completely.
Does anyone know what we can do to finally stop it?
I have no clients, and it still does not want to stop rank1. Funny thing is, while trying to fix this by restarting mdses, I sometimes see a list of clients popping up in the dashboard, even though no clients are connected..
Configuring debuglogging shows me the following: https://apac01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fp.6core.n...
I think your case is different from mine. Your logs show “waiting for stray to migrate”. I didn’t see this.
I have quite a lot of hardlinks on this filesystem, which I've seen issue with 'No space left on device'. I have mds_bal_fragment_size_max set to 200000 to mitigate that.
The message 'waiting for strays to migrate' makes me feel like I should push the MDS to migrate them somehow .. But how?
-- Mark Schouten | Tuxis B.V. KvK: 74698818 | https://apac01.safelinks.protection.outlook.com/?url=http%3A%2F%2Fwww.tuxis.... T: +31 318 200208 | info@tuxis.nl _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (3)
-
Mark Schouten
-
Martin Rasmus Lundquist Hansen
-
胡 玮文