New milestones for reef and quincy
Hello Dev Leads - just a friendly reminder that we've defined new milestones for reef and quincy releases: 18.2.2 and 17.2.8 Please assign the PRs that are to be tested and merged as part of those milestones and add the "needs-qa" label. TIA
Am 17.01.24 um 17:34 schrieb Yuri Weinstein:
Dev Leads - just a friendly reminder that we've defined new milestones for reef and quincy releases: 18.2.2 and 17.2.8
Please assign the PRs that are to be tested and merged as part of those milestones and add the "needs-qa" label.
After being hit repeatedly at several customer sites, we would appreciate these bugs being fixed in one of those releases: https://tracker.ceph.com/issues/63364 MDS_CLIENT_OLDEST_TID: 15 clients failing to advance oldest client/flush tid https://tracker.ceph.com/issues/62702 MDS slow requests for the internal 'rename' requests https://tracker.ceph.com/issues/62257 mds: blocklist clients that are not advancing `oldest_client_tid` Thanks for your work! Amon Ott -- Dr. Amon Ott m-privacy GmbH Tel: +49 30 24342334 Werner-Voß-Damm 62 Fax: +49 30 99296856 12101 Berlin http://www.m-privacy.de Amtsgericht Charlottenburg, HRB 84946 Geschäftsführer: Dipl.-Kfm. Holger Maczkowsky, Roman Maczkowsky GnuPG-Key-ID: 0x2DD3A649
Hi Amon, The trackers you mention would be a part of the next quincy and reef releases.
Hi Venky! Am 23.01.24 um 05:58 schrieb Venky Shankar:
The trackers you mention would be a part of the next quincy and reef releases.
Thank you! For now, the systems seem to run fine with Linux kernel 5.10 on the client side, so we have a workaround. The problems started after we changed to 6.1 - maybe that info helps identifying what triggered these bugs. Amon Ott -- Dr. Amon Ott m-privacy GmbH Tel: +49 30 24342334 Werner-Voß-Damm 62 Fax: +49 30 99296856 12101 Berlin http://www.m-privacy.de Amtsgericht Charlottenburg, HRB 84946 Geschäftsführer: Dipl.-Kfm. Holger Maczkowsky, Roman Maczkowsky GnuPG-Key-ID: 0x2DD3A649
On Tue, Jan 23, 2024 at 7:59 AM Amon Ott <lists@compuniverse.de> wrote:
Hi Venky!
Am 23.01.24 um 05:58 schrieb Venky Shankar:
The trackers you mention would be a part of the next quincy and reef releases.
Thank you! For now, the systems seem to run fine with Linux kernel 5.10 on the client side, so we have a workaround. The problems started after we changed to 6.1 - maybe that info helps identifying what triggered these bugs.
Hi Amon, Is any of your CephFS pools nearfull? If so, you are likely affected by [1] and would need to add additional capacity or bump the nearfull threshold. [1] https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/E2QJ6K4UBYFM... Thanks, Ilya
Am 23.01.24 um 13:09 schrieb Ilya Dryomov:
On Tue, Jan 23, 2024 at 7:59 AM Amon Ott <lists@compuniverse.de> wrote:
Am 23.01.24 um 05:58 schrieb Venky Shankar:
The trackers you mention would be a part of the next quincy and reef releases.
Thank you! For now, the systems seem to run fine with Linux kernel 5.10 on the client side, so we have a workaround. The problems started after we changed to 6.1 - maybe that info helps identifying what triggered these bugs.
Is any of your CephFS pools nearfull? If so, you are likely affected by [1] and would need to add additional capacity or bump the nearfull threshold.
[1] https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/E2QJ6K4UBYFM...
Thank you, but no, the CephFS had plenty of space and we had several with these issues. The clients do not respond to capability release requests in time or at all. We see lots of slow requests in Ceph log, sometimes MDS goes readonly. In both cases the whole cluster stalls. As workarounds, we started avoiding renames and moved some temp file areas to local disks. It helped, but was not enough. Unfortunately, though blocking a client helps with slow requests, it hangs that client and it needs a hard reboot. Mounting with recover_session=clean did not help. With kernel 5.10 on the clients everything works fine. Amon.
On 17-01-2024 17:34, Yuri Weinstein wrote:
Hello
Dev Leads - just a friendly reminder that we've defined new milestones for reef and quincy releases: 18.2.2 and 17.2.8
Please assign the PRs that are to be tested and merged as part of those milestones and add the "needs-qa" label.
Here a list of PRs that I would like to see in 18.2.2 https://github.com/ceph/ceph/pull/55335 / https://tracker.ceph.com/issues/64197 <- still needs-qa label This will make 18.2.2 the fastest release to date for all flash clusters with OSDs encrypted with dmcrypt. Tested by us on 16.2.11 / 18.2.0. https://github.com/ceph/ceph/pull/53250 (perfcount for bluestore/bluefs allocator), already has "needs-qa" label For Ceph clusters that suffer from fragmented OSDs, this will be highly valuable (we used a backport from this PR in 16.2.11 to identify performance degradation). It will also be useful to see the impact of the already backported PR 54434 (https://github.com/ceph/ceph/pull/54434). https://github.com/ceph/ceph/pull/54285 / https://tracker.ceph.com/issues/63448 (cephadm discovery fails IPv6 only) <- still needs-qa label Thanks, Gr. Stefan
participants (5)
-
Amon Ott
-
Ilya Dryomov
-
Stefan Kooman
-
Venky Shankar
-
Yuri Weinstein