Ceph FS not releasing space after file deletion
Dear ceph users, we are using CEPH since Jewel and now (with Mimic) are facing a problem that is impeding us to use it appropriately. After deleting some big backup files, we have noticed that their space has not returned to the pool's free space. Considering the mentioned pool, du shows that we have a total of 247TB allocated in files and "ceph df" says we have 293TB of used space. How can we debug this problem and force ceph to release unallocated space? Gustavo.
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]). The Dennis Kramer's case [3] led me to believe that I need to recreate the FS. But I refuse to (dis)believe that CEPH hasn't a repair tool for it. I thought "cephfs-table-tool take_inos" could be the answer for my problem, but the message [4] were not clear enough. Can I run the command without resetting the inodes? [1] Error at ceph -w - https://pastebin.com/imNqBdmH [2] Error at mds.log - https://pastebin.com/rznkzLHG [3] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027845.h... [4] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027935.h...
On Tue, Sep 3, 2019 at 3:39 PM Guilherme <guilherme.geronimo@gmail.com> wrote:
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]).
Might be related to this ticket: https://tracker.ceph.com/issues/38452 -- Patrick Donnelly, Ph.D. He / Him / His Senior Software Engineer Red Hat Sunnyvale, CA GPG: 19F28A586F808C2402351B93C3301A3E258DD79D
On Wed, Sep 4, 2019 at 6:39 AM Guilherme <guilherme.geronimo@gmail.com> wrote:
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]).
The Dennis Kramer's case [3] led me to believe that I need to recreate the FS. But I refuse to (dis)believe that CEPH hasn't a repair tool for it. I thought "cephfs-table-tool take_inos" could be the answer for my problem, but the message [4] were not clear enough. Can I run the command without resetting the inodes?
[1] Error at ceph -w - https://pastebin.com/imNqBdmH [2] Error at mds.log - https://pastebin.com/rznkzLHG
For the mds crash issue. 'cephfs-data-scan scan_link' of nautilus version (14.2.2) should fix it. snaptable. You don't need to upgrade whole cluster. Just install nautilus in a temp machine or compile ceph from source.
[3] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027845.h... [4] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027935.h... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Thank you, Yan. It took like 10 minutes to execute the scan_links. I believe the number of Lost+Found decreased in 60%, but the rest of them are still causing the MDS crash. Any other suggestion? =D []'s Arthur (aKa Guilherme Geronimo) On 10/09/2019 23:51, Yan, Zheng wrote:
On Wed, Sep 4, 2019 at 6:39 AM Guilherme <guilherme.geronimo@gmail.com> wrote:
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]).
The Dennis Kramer's case [3] led me to believe that I need to recreate the FS. But I refuse to (dis)believe that CEPH hasn't a repair tool for it. I thought "cephfs-table-tool take_inos" could be the answer for my problem, but the message [4] were not clear enough. Can I run the command without resetting the inodes?
[1] Error at ceph -w - https://pastebin.com/imNqBdmH [2] Error at mds.log - https://pastebin.com/rznkzLHG For the mds crash issue. 'cephfs-data-scan scan_link' of nautilus version (14.2.2) should fix it. snaptable. You don't need to upgrade whole cluster. Just install nautilus in a temp machine or compile ceph from source.
[3] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027845.h... [4] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027935.h... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
please send me crash log On Tue, Sep 17, 2019 at 12:56 AM Guilherme Geronimo <guilherme.geronimo@gmail.com> wrote:
Thank you, Yan.
It took like 10 minutes to execute the scan_links. I believe the number of Lost+Found decreased in 60%, but the rest of them are still causing the MDS crash.
Any other suggestion?
=D
[]'s Arthur (aKa Guilherme Geronimo)
On 10/09/2019 23:51, Yan, Zheng wrote:
On Wed, Sep 4, 2019 at 6:39 AM Guilherme <guilherme.geronimo@gmail.com> wrote:
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]).
The Dennis Kramer's case [3] led me to believe that I need to recreate the FS. But I refuse to (dis)believe that CEPH hasn't a repair tool for it. I thought "cephfs-table-tool take_inos" could be the answer for my problem, but the message [4] were not clear enough. Can I run the command without resetting the inodes?
[1] Error at ceph -w - https://pastebin.com/imNqBdmH [2] Error at mds.log - https://pastebin.com/rznkzLHG For the mds crash issue. 'cephfs-data-scan scan_link' of nautilus version (14.2.2) should fix it. snaptable. You don't need to upgrade whole cluster. Just install nautilus in a temp machine or compile ceph from source.
[3] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027845.h... [4] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027935.h... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Here it is: https://pastebin.com/SAsqnWDi The command: timeout 10 rm /mnt/ceph/lost+found/100002430c8 ; umount -f /mnt/ceph On 17/09/2019 00:51, Yan, Zheng wrote:
please send me crash log
On Tue, Sep 17, 2019 at 12:56 AM Guilherme Geronimo <guilherme.geronimo@gmail.com> wrote:
Thank you, Yan.
It took like 10 minutes to execute the scan_links. I believe the number of Lost+Found decreased in 60%, but the rest of them are still causing the MDS crash.
Any other suggestion?
=D
[]'s Arthur (aKa Guilherme Geronimo)
On 10/09/2019 23:51, Yan, Zheng wrote:
On Wed, Sep 4, 2019 at 6:39 AM Guilherme <guilherme.geronimo@gmail.com> wrote:
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]).
The Dennis Kramer's case [3] led me to believe that I need to recreate the FS. But I refuse to (dis)believe that CEPH hasn't a repair tool for it. I thought "cephfs-table-tool take_inos" could be the answer for my problem, but the message [4] were not clear enough. Can I run the command without resetting the inodes?
[1] Error at ceph -w - https://pastebin.com/imNqBdmH [2] Error at mds.log - https://pastebin.com/rznkzLHG For the mds crash issue. 'cephfs-data-scan scan_link' of nautilus version (14.2.2) should fix it. snaptable. You don't need to upgrade whole cluster. Just install nautilus in a temp machine or compile ceph from source.
[3] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027845.h... [4] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027935.h... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Btw: root@deployer:~# cephfs-data-scan -v ceph version 14.2.4 (75f4de193b3ea58512f204623e6c5a16e6c1e1ba) nautilus (stable) On 19/09/2019 13:38, Guilherme Geronimo wrote:
Here it is: https://pastebin.com/SAsqnWDi
The command:
timeout 10 rm /mnt/ceph/lost+found/100002430c8 ; umount -f /mnt/ceph
On 17/09/2019 00:51, Yan, Zheng wrote:
please send me crash log
On Tue, Sep 17, 2019 at 12:56 AM Guilherme Geronimo <guilherme.geronimo@gmail.com> wrote:
Thank you, Yan.
It took like 10 minutes to execute the scan_links. I believe the number of Lost+Found decreased in 60%, but the rest of them are still causing the MDS crash.
Any other suggestion?
=D
[]'s Arthur (aKa Guilherme Geronimo)
On 10/09/2019 23:51, Yan, Zheng wrote:
On Wed, Sep 4, 2019 at 6:39 AM Guilherme <guilherme.geronimo@gmail.com> wrote:
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]).
The Dennis Kramer's case [3] led me to believe that I need to recreate the FS. But I refuse to (dis)believe that CEPH hasn't a repair tool for it. I thought "cephfs-table-tool take_inos" could be the answer for my problem, but the message [4] were not clear enough. Can I run the command without resetting the inodes?
[1] Error at ceph -w - https://pastebin.com/imNqBdmH [2] Error at mds.log - https://pastebin.com/rznkzLHG For the mds crash issue. 'cephfs-data-scan scan_link' of nautilus version (14.2.2) should fix it. snaptable. You don't need to upgrade whole cluster. Just install nautilus in a temp machine or compile ceph from source.
[3] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027845.h... [4] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027935.h... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On Fri, Sep 20, 2019 at 12:38 AM Guilherme Geronimo <guilherme.geronimo@gmail.com> wrote:
Here it is: https://pastebin.com/SAsqnWDi
please set debug_mds to 10 and send detailed log to me
The command:
timeout 10 rm /mnt/ceph/lost+found/100002430c8 ; umount -f /mnt/ceph
On 17/09/2019 00:51, Yan, Zheng wrote:
please send me crash log
On Tue, Sep 17, 2019 at 12:56 AM Guilherme Geronimo <guilherme.geronimo@gmail.com> wrote:
Thank you, Yan.
It took like 10 minutes to execute the scan_links. I believe the number of Lost+Found decreased in 60%, but the rest of them are still causing the MDS crash.
Any other suggestion?
=D
[]'s Arthur (aKa Guilherme Geronimo)
On 10/09/2019 23:51, Yan, Zheng wrote:
On Wed, Sep 4, 2019 at 6:39 AM Guilherme <guilherme.geronimo@gmail.com> wrote:
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]).
The Dennis Kramer's case [3] led me to believe that I need to recreate the FS. But I refuse to (dis)believe that CEPH hasn't a repair tool for it. I thought "cephfs-table-tool take_inos" could be the answer for my problem, but the message [4] were not clear enough. Can I run the command without resetting the inodes?
[1] Error at ceph -w - https://pastebin.com/imNqBdmH [2] Error at mds.log - https://pastebin.com/rznkzLHG For the mds crash issue. 'cephfs-data-scan scan_link' of nautilus version (14.2.2) should fix it. snaptable. You don't need to upgrade whole cluster. Just install nautilus in a temp machine or compile ceph from source.
[3] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027845.h... [4] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027935.h... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Wow, something is seriously wrong: if I turn on the "debug mds = 10" (on [mds] context) and restart MDS, it instantaneously crashes! If I comment it, everything is alright. https://pastebin.com/GPcKRmR9 On 19/09/2019 22:34, Yan, Zheng wrote:
On Fri, Sep 20, 2019 at 12:38 AM Guilherme Geronimo <guilherme.geronimo@gmail.com> wrote:
Here it is: https://pastebin.com/SAsqnWDi
please set debug_mds to 10 and send detailed log to me
The command:
timeout 10 rm /mnt/ceph/lost+found/100002430c8 ; umount -f /mnt/ceph
On 17/09/2019 00:51, Yan, Zheng wrote:
please send me crash log
On Tue, Sep 17, 2019 at 12:56 AM Guilherme Geronimo <guilherme.geronimo@gmail.com> wrote:
Thank you, Yan.
It took like 10 minutes to execute the scan_links. I believe the number of Lost+Found decreased in 60%, but the rest of them are still causing the MDS crash.
Any other suggestion?
=D
[]'s Arthur (aKa Guilherme Geronimo)
On 10/09/2019 23:51, Yan, Zheng wrote:
On Wed, Sep 4, 2019 at 6:39 AM Guilherme <guilherme.geronimo@gmail.com> wrote:
Dear CEPHers, Adding some comments to my colleague's post: we are running Mimic 13.2.6 and struggling with 2 issues (that might be related): 1) After a "lack of space" event we've tried to remove a 40TB file. The file is not there anymore, but no space was released. No process is using the file either. 2) There are many files in /lost+found (~25TB|~5%). Every time we try to remove a file, MDS crashes ([1,2]).
The Dennis Kramer's case [3] led me to believe that I need to recreate the FS. But I refuse to (dis)believe that CEPH hasn't a repair tool for it. I thought "cephfs-table-tool take_inos" could be the answer for my problem, but the message [4] were not clear enough. Can I run the command without resetting the inodes?
[1] Error at ceph -w - https://pastebin.com/imNqBdmH [2] Error at mds.log - https://pastebin.com/rznkzLHG For the mds crash issue. 'cephfs-data-scan scan_link' of nautilus version (14.2.2) should fix it. snaptable. You don't need to upgrade whole cluster. Just install nautilus in a temp machine or compile ceph from source.
[3] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027845.h... [4] Discussion - http://lists.opennebula.org/pipermail/ceph-users-ceph.com/2018-July/027935.h... _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (5)
-
Guilherme
-
Guilherme Geronimo
-
Gustavo Tonini
-
Patrick Donnelly
-
Yan, Zheng