issue with fuse client on multifs auth PR
Hi all, I am testing my multiFS auth caps PR[1] using teuthology. I have been seeing this error[2] in logs. I thought the issue was genuinely that directory was absent so I wrote a patch to create directory when absent[3] and ran the tests again which fixes the issue but the test jobs hits another error[4]. Also, it turned out that the patch[3] was not the correct fix (the directory "/sys/fs/fuse/connections" should've been created with the mount command) from the conversation here[5][6][7]. So I am certainly not on the right track here. Any suggestions/ideas on why do I get error on [2] or on how do I investigate further into the reason for this error? [8][9] are teuthology logs where I get error on [2] and [10][11] are teuthology logs where I get error[4]. Thanks, - Rishabh [1] https://github.com/ceph/ceph/pull/32581 [2] https://gist.github.com/rishabh-d-dave/eef6cdb21f54a95edec25d412e52d09e#file... [3] https://github.com/ceph/ceph/pull/34839/commits/17db08fedb9f24af3b874756b769... [4] https://gist.github.com/rishabh-d-dave/eef6cdb21f54a95edec25d412e52d09e#file... [5] https://github.com/ceph/ceph/pull/34839#pullrequestreview-405003346 [6] https://github.com/ceph/ceph/pull/34839#issuecomment-623906416 [7] https://github.com/ceph/ceph/pull/34839#issuecomment-625171252 [8] http://qa-proxy.ceph.com/teuthology/rishabh-2020-05-11_05:57:36-fs-wip-risha... [9] [10] http://qa-proxy.ceph.com/teuthology/rishabh-2020-05-08_15:19:02-fs-wip-risha... [11] http://qa-proxy.ceph.com/teuthology/rishabh-2020-05-08_13:02:19-fs-wip-risha...
This job only takes about 15 minutes to get to the point where it errors but then runs for several hours after that. I'd suggest running it again and inspecting the status of the daemons once you know the error has occurred. On Tue, May 12, 2020 at 1:18 PM Rishabh Dave <ridave@redhat.com> wrote:
Hi all,
I am testing my multiFS auth caps PR[1] using teuthology. I have been seeing this error[2] in logs. I thought the issue was genuinely that directory was absent so I wrote a patch to create directory when absent[3] and ran the tests again which fixes the issue but the test jobs hits another error[4].
Also, it turned out that the patch[3] was not the correct fix (the directory "/sys/fs/fuse/connections" should've been created with the mount command) from the conversation here[5][6][7]. So I am certainly not on the right track here. Any suggestions/ideas on why do I get error on [2] or on how do I investigate further into the reason for this error?
[8][9] are teuthology logs where I get error on [2] and [10][11] are teuthology logs where I get error[4].
Thanks, - Rishabh
[1] https://github.com/ceph/ceph/pull/32581 [2] https://gist.github.com/rishabh-d-dave/eef6cdb21f54a95edec25d412e52d09e#file... [3] https://github.com/ceph/ceph/pull/34839/commits/17db08fedb9f24af3b874756b769... [4] https://gist.github.com/rishabh-d-dave/eef6cdb21f54a95edec25d412e52d09e#file... [5] https://github.com/ceph/ceph/pull/34839#pullrequestreview-405003346 [6] https://github.com/ceph/ceph/pull/34839#issuecomment-623906416 [7] https://github.com/ceph/ceph/pull/34839#issuecomment-625171252 [8] http://qa-proxy.ceph.com/teuthology/rishabh-2020-05-11_05:57:36-fs-wip-risha... [9] [10] http://qa-proxy.ceph.com/teuthology/rishabh-2020-05-08_15:19:02-fs-wip-risha... [11] http://qa-proxy.ceph.com/teuthology/rishabh-2020-05-08_13:02:19-fs-wip-risha... _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
-- Cheers, Brad
On Tue, 12 May 2020 at 10:47, Brad Hubbard <bhubbard@redhat.com> wrote:
This job only takes about 15 minutes to get to the point where it errors but then runs for several hours after that. I'd suggest running it again and inspecting the status of the daemons once you know the error has occurred.
I logged into the machine and checked for existence of /sys/fs/fuse/connection and the mountpoint and checked whether Ceph FS was mounted and finally, checked Ceph cluster's status. The results[1][2][3] were positive for all the checks, So I figured that the execution probably needs to wait a bit before running "ls /sys/fs/fuse/connections"[4] and it worked. The execution moved past that point and testsuite crashed at a different point. I am trying to find out exact cause, I'll mail on this thread in case I can't. Thanks for the help, Brad! [1] https://gist.github.com/rishabh-d-dave/eef6cdb21f54a95edec25d412e52d09e#file... [2] https://gist.github.com/rishabh-d-dave/eef6cdb21f54a95edec25d412e52d09e#file... [3] https://gist.github.com/rishabh-d-dave/eef6cdb21f54a95edec25d412e52d09e#file... [4] https://github.com/rishabh-d-dave/ceph/commit/325c7f0447112a90dea656c5296852...; see the commit with title "DNM: let's wait before checking connection dir" in case commit SHA changes
I still see the error at [1] but I don't know the exact cause. Since this error occurs only on RHEL and CentOS and not on Ubuntu (see [2] and [3]), I suspect either I changed some code in qa/tasks/cephfs that was specific to these distros or my code exposes a bug in ceph-fuse on these distros. [1] https://gist.github.com/rishabh-d-dave/eef6cdb21f54a95edec25d412e52d09e#file... [2] http://pulpito.ceph.com/rishabh-2020-05-14_09:40:39-fs-wip-rishabh-15070-dis... [3] http://pulpito.ceph.com/rishabh-2020-05-14_07:41:09-fs-wip-rishabh-15070-dis...
participants (2)
-
Brad Hubbard
-
Rishabh Dave