21 Jul
2023
21 Jul
'23
1:58 p.m.
At 01:27 this morning I received the first email about MDS cache is too large (mailing happens every 15 minutes if something happens). Looking into it, it was again a standby-replay host which stops working.
At 01:00 a few rsync processes start in parallel on a client machine. This copies data from a NFS share to Cephfs share to sync the latest changes. (we want to switch to Cephfs in the near future).
This crashing of the standby-replay mds happend a couple times now, so I think it would be good to get some help. Where should I look next?
What do you mean with crashing? Is the container just getting OOM and killed and restarted? Then you just have to adapt your settings not?