Heads up -- 18.2.6 regression
Hi all, Just a quick heads up -- 18.2.6 has a bluefs regression [1] which was not present in 18.2.4. Please avoid upgrading until further notice. Regards, Dan [1] Introduced in https://tracker.ceph.com/issues/65356 and fixed in https://tracker.ceph.com/issues/69764 -- Dan van der Ster Ceph Executive Council | CTO @ CLYSO https://clyso.com | dan.vanderster@clyso.com
Hi Dan, bad luck, we just upgraded today! Is it so severe that we should plan to downgrade to 18.2.4 if feasible? Or if the issue is not present in Squid, to upgrade to 19.2.2? Best regards, Michel Sent from my mobile Le 30 avril 2025 19:32:45 Dan van der Ster <dan.vanderster@clyso.com> a écrit :
Hi all,
Just a quick heads up -- 18.2.6 has a bluefs regression [1] which was not present in 18.2.4. Please avoid upgrading until further notice.
Regards, Dan
[1] Introduced in https://tracker.ceph.com/issues/65356 and fixed in https://tracker.ceph.com/issues/69764
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO https://clyso.com | dan.vanderster@clyso.com _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Michel, 19.2.2 should be immune from this bug, but I don't know if 19.2.2 has enough real life usage to check for other potential issues. We're working on an 18.2.7 hotfix now -- my personal suggestion is to wait for that. Cheers, Dan On Wed, Apr 30, 2025 at 11:51 AM Michel Jouvin <michel.jouvin@ijclab.in2p3.fr> wrote:
Hi Dan,
bad luck, we just upgraded today! Is it so severe that we should plan to downgrade to 18.2.4 if feasible? Or if the issue is not present in Squid, to upgrade to 19.2.2?
Best regards,
Michel Sent from my mobile
Le 30 avril 2025 19:32:45 Dan van der Ster <dan.vanderster@clyso.com> a écrit :
Hi all,
Just a quick heads up -- 18.2.6 has a bluefs regression [1] which was not present in 18.2.4. Please avoid upgrading until further notice.
Regards, Dan
[1] Introduced in https://tracker.ceph.com/issues/65356 and fixed in https://tracker.ceph.com/issues/69764
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO https://clyso.com | dan.vanderster@clyso.com _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO https://clyso.com | dan.vanderster@clyso.com
Hi Dan, Thanks for the advice. Is it possible to get a bit more information on what is the risk of this bug? I tried to read the issues you referenced and the other referenced in the Issues but I didn't get the full picture... Is it just an osd crash or is there a risk of data corruption? I'm afraid it is more the last one... Cheers, Michel Sent from my mobile Le 30 avril 2025 21:14:29 Dan van der Ster <dan.vanderster@clyso.com> a écrit :
Hi Michel,
19.2.2 should be immune from this bug, but I don't know if 19.2.2 has enough real life usage to check for other potential issues.
We're working on an 18.2.7 hotfix now -- my personal suggestion is to wait for that.
Cheers, Dan
On Wed, Apr 30, 2025 at 11:51 AM Michel Jouvin <michel.jouvin@ijclab.in2p3.fr> wrote:
Hi Dan,
bad luck, we just upgraded today! Is it so severe that we should plan to downgrade to 18.2.4 if feasible? Or if the issue is not present in Squid, to upgrade to 19.2.2?
Best regards,
Michel Sent from my mobile
Le 30 avril 2025 19:32:45 Dan van der Ster <dan.vanderster@clyso.com> a écrit :
Hi all,
Just a quick heads up -- 18.2.6 has a bluefs regression [1] which was not present in 18.2.4. Please avoid upgrading until further notice.
Regards, Dan
[1] Introduced in https://tracker.ceph.com/issues/65356 and fixed in https://tracker.ceph.com/issues/69764
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO https://clyso.com | dan.vanderster@clyso.com _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO https://clyso.com | dan.vanderster@clyso.com
Hi Michel, Our understanding is that the bug is quite rare -- it seems to me that there is something about the workload / IO pattern that triggers it. I raised the alarm about this because the first cluster I upgraded to 18.2.6 had 4 OSDs crash immediately during upgrade, luckily all in the same failure domain. (So for that cluster, we paused the upgrade and are awaiting 18.2.7 to resume). 17.2.8 has the bug too -- there are reports about that in https://tray cker.ceph.com/issues/69764. My overall guess/hope -- if an OSD didn't hit the crash yet, then it's hopefully going to survive until we get 18.2.7 out. Cheers, Dan On Thu, May 1, 2025 at 12:17 AM Michel Jouvin <michel.jouvin@ijclab.in2p3.fr> wrote:
Hi Dan,
Thanks for the advice. Is it possible to get a bit more information on what is the risk of this bug? I tried to read the issues you referenced and the other referenced in the Issues but I didn't get the full picture... Is it just an osd crash or is there a risk of data corruption? I'm afraid it is more the last one...
Cheers,
Michel Sent from my mobile
Le 30 avril 2025 21:14:29 Dan van der Ster <dan.vanderster@clyso.com> a écrit :
Hi Michel,
19.2.2 should be immune from this bug, but I don't know if 19.2.2 has enough real life usage to check for other potential issues.
We're working on an 18.2.7 hotfix now -- my personal suggestion is to wait for that.
Cheers, Dan
On Wed, Apr 30, 2025 at 11:51 AM Michel Jouvin <michel.jouvin@ijclab.in2p3.fr> wrote:
Hi Dan,
bad luck, we just upgraded today! Is it so severe that we should plan to downgrade to 18.2.4 if feasible? Or if the issue is not present in Squid, to upgrade to 19.2.2?
Best regards,
Michel Sent from my mobile
Le 30 avril 2025 19:32:45 Dan van der Ster <dan.vanderster@clyso.com> a écrit :
Hi all,
Just a quick heads up -- 18.2.6 has a bluefs regression [1] which was not present in 18.2.4. Please avoid upgrading until further notice.
Regards, Dan
[1] Introduced in https://tracker.ceph.com/issues/65356 and fixed in https://tracker.ceph.com/issues/69764
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO https://clyso.com | dan.vanderster@clyso.com _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO https://clyso.com | dan.vanderster@clyso.com
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO Try our Ceph Analyzer -- https://analyzer.clyso.com/ https://clyso.com | dan.vanderster@clyso.com
Thanks for the notice Dan. We just finished upgrading to Red Hat Ceph Storage 6 and the next phase is for us to upgrade to 7. Is 18.2.6 Red Hat Ceph Storage 7.x? In that case I'll be sure to not upgrade until we get the "green light" from you. I assume the fix is on the scale of weeks out, not months? Thanks again!
Hi Alex, I'm not sure what exactly is in RHCS 7.x. I think you can `rpm -q --changelog` to see the exact patches -- the commits to look for are: -- the bug: os/bluestore: fix bluefs log runway enospc -- the fix: os/bluestore: fix _extend_log seq advance Cheers, Dan On Thu, May 1, 2025 at 8:42 AM Alex <mr.alexey@gmail.com> wrote:
Thanks for the notice Dan.
We just finished upgrading to Red Hat Ceph Storage 6 and the next phase is for us to upgrade to 7. Is 18.2.6 Red Hat Ceph Storage 7.x?
In that case I'll be sure to not upgrade until we get the "green light" from you. I assume the fix is on the scale of weeks out, not months?
Thanks again!
-- Dan van der Ster Ceph Executive Council | CTO @ CLYSO Try our Ceph Analyzer -- https://analyzer.clyso.com/ https://clyso.com | dan.vanderster@clyso.com
Thanks. According to Red Hat Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid Is the bug in Reef or Squid?
From a previous message by Dan, Reef only. Michel Sent from my mobile Le 1 mai 2025 18:51:52 Alex <mr.alexey@gmail.com> a écrit :
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I was too quick... Reef >= 18.2.5. so mid you are lucky RH Ceph 7 is <= 18.2.4 and immune. Michel Sent from my mobile Le 1 mai 2025 19:18:51 Michel Jouvin <michel.jouvin@ijclab.in2p3.fr> a écrit :
From a previous message by Dan, Reef only.
Michel Sent from my mobile Le 1 mai 2025 18:51:52 Alex <mr.alexey@gmail.com> a écrit :
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Michel, RHCS 7 is based on v18.2.1 but includes backported features cherry-picked from newer releases. I reviewed the release notes through latest 7.1z3 and found no references to the bug introduced by https://tracker.ceph.com/issues/65356. Can't definitively confirm the absence of this bug in latest 7.1z3 (since RH BZs don't always link to upstream trackers), but none of the BZs descriptions appear to relate to this particular issue. Cheers, Frédéric. ----- Le 1 Mai 25, à 19:21, Michel Jouvin michel.jouvin@ijclab.in2p3.fr a écrit :
I was too quick... Reef >= 18.2.5. so mid you are lucky RH Ceph 7 is <= 18.2.4 and immune.
Michel Sent from my mobile Le 1 mai 2025 19:18:51 Michel Jouvin <michel.jouvin@ijclab.in2p3.fr> a écrit :
From a previous message by Dan, Reef only.
Michel Sent from my mobile Le 1 mai 2025 18:51:52 Alex <mr.alexey@gmail.com> a écrit :
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
These are my notes (come with no warranty) about affected / immune versions as far as upstream Ceph is concerned: Quincy: 17.2.7: OK – does not `void BlueFS::_extend_log` at all 17.2.8: BUGGY (missing https://github.com/ceph/ceph/pull/53732/) Reef: 18.2.4: OK – does not `void BlueFS::_extend_log` at all 18.2.5 (branch as 09/04): BUGGY (missing https://github.com/ceph/ceph/pull/53732/) 18.2.6 (branch as 06/05): BUGGY (missing https://github.com/ceph/ceph/pull/53732/) Squid: 19.2.1: OK (has https://github.com/ceph/ceph/pull/53732/) 19.2.2 (branch as of 09/04): OK (has https://github.com/ceph/ceph/pull/53732/) Main (as 09/04): OK (has https://github.com/ceph/ceph/pull/53732/) Relevant backports: - Quincy: https://github.com/ceph/ceph/pull/63122 (<-- this is very recent. 53732 applies cleanly on top of v17.2.8) - Reef: https://github.com/ceph/ceph/pull/61653 Cheers, Enrico On 5/2/25 09:46, Frédéric Nass wrote:
Hi Michel,
RHCS 7 is based on v18.2.1 but includes backported features cherry-picked from newer releases. I reviewed the release notes through latest 7.1z3 and found no references to the bug introduced by https://tracker.ceph.com/issues/65356. Can't definitively confirm the absence of this bug in latest 7.1z3 (since RH BZs don't always link to upstream trackers), but none of the BZs descriptions appear to relate to this particular issue.
Cheers, Frédéric.
----- Le 1 Mai 25, à 19:21, Michel Jouvin michel.jouvin@ijclab.in2p3.fr a écrit :
I was too quick... Reef >= 18.2.5. so mid you are lucky RH Ceph 7 is <= 18.2.4 and immune.
Michel Sent from my mobile Le 1 mai 2025 19:18:51 Michel Jouvin <michel.jouvin@ijclab.in2p3.fr> a écrit :
From a previous message by Dan, Reef only.
Michel Sent from my mobile Le 1 mai 2025 18:51:52 Alex <mr.alexey@gmail.com> a écrit :
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Enrico Bocchi CERN European Laboratory for Particle Physics IT - Storage & Data Management - General Storage Services Mailbox: G20500 - Office: 31-2-010 1211 Genève 23 Switzerland
Hello Alex As per Dan bug is in Reef which is v18.2.6 in open ceph. I also upgraded my cluster to 18.2.6 before I saw the first message of this mail chain and I am having 120osds but yet I have not seen any issue where’s as one of my host remain down for 24hrs with 24 osds on it and now I joined back and all osds came active. Regards Dev On Thu, 1 May 2025 at 9:51 AM, Alex <mr.alexey@gmail.com> wrote:
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
One of the referenced issues seems to indicate it may be in quincy too? does anyone know if thats true, and when it was introduced? Does it go back even farther? Thanks, Kevin ________________________________________ From: Devender Singh <devender@netskrt.io> Sent: Thursday, May 1, 2025 11:40 AM To: Alex Cc: Dan van der Ster; ceph-users Subject: [ceph-users] Re: Heads up -- 18.2.6 regression Check twice before you click! This email originated from outside PNNL. Hello Alex As per Dan bug is in Reef which is v18.2.6 in open ceph. I also upgraded my cluster to 18.2.6 before I saw the first message of this mail chain and I am having 120osds but yet I have not seen any issue where’s as one of my host remain down for 24hrs with 24 osds on it and now I joined back and all osds came active. Regards Dev On Thu, 1 May 2025 at 9:51 AM, Alex <mr.alexey@gmail.com> wrote:
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
In Slack it was confirmed for 17.2.8, other versions have not been mentioned (yet). Zitat von "Fox, Kevin M" <Kevin.Fox@pnnl.gov>:
One of the referenced issues seems to indicate it may be in quincy too? does anyone know if thats true, and when it was introduced? Does it go back even farther?
Thanks, Kevin
________________________________________ From: Devender Singh <devender@netskrt.io> Sent: Thursday, May 1, 2025 11:40 AM To: Alex Cc: Dan van der Ster; ceph-users Subject: [ceph-users] Re: Heads up -- 18.2.6 regression
Check twice before you click! This email originated from outside PNNL.
Hello Alex
As per Dan bug is in Reef which is v18.2.6 in open ceph.
I also upgraded my cluster to 18.2.6 before I saw the first message of this mail chain and I am having 120osds but yet I have not seen any issue where’s as one of my host remain down for 24hrs with 24 osds on it and now I joined back and all osds came active.
Regards Dev
On Thu, 1 May 2025 at 9:51 AM, Alex <mr.alexey@gmail.com> wrote:
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Kevin, You are right, Quincy is also affected in 17.2.8 but as the question was about Reef and Squid I didn't mention it. If you read the issues mentioned by Dan, you will see that it seems there is no plan to backport the patch to Quincy as it is EOL. Michel Le 01/05/2025 à 21:51, Eugen Block a écrit :
In Slack it was confirmed for 17.2.8, other versions have not been mentioned (yet).
Zitat von "Fox, Kevin M" <Kevin.Fox@pnnl.gov>:
One of the referenced issues seems to indicate it may be in quincy too? does anyone know if thats true, and when it was introduced? Does it go back even farther?
Thanks, Kevin
________________________________________ From: Devender Singh <devender@netskrt.io> Sent: Thursday, May 1, 2025 11:40 AM To: Alex Cc: Dan van der Ster; ceph-users Subject: [ceph-users] Re: Heads up -- 18.2.6 regression
Check twice before you click! This email originated from outside PNNL.
Hello Alex
As per Dan bug is in Reef which is v18.2.6 in open ceph.
I also upgraded my cluster to 18.2.6 before I saw the first message of this mail chain and I am having 120osds but yet I have not seen any issue where’s as one of my host remain down for 24hrs with 24 osds on it and now I joined back and all osds came active.
Regards Dev
On Thu, 1 May 2025 at 9:51 AM, Alex <mr.alexey@gmail.com> wrote:
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
As Igor mentioned in slack, it has been proposed to CSC to make a hot fix for Quincy as well (17.2.9). Zitat von Michel Jouvin <michel.jouvin@ijclab.in2p3.fr>:
Kevin,
You are right, Quincy is also affected in 17.2.8 but as the question was about Reef and Squid I didn't mention it. If you read the issues mentioned by Dan, you will see that it seems there is no plan to backport the patch to Quincy as it is EOL.
Michel
Le 01/05/2025 à 21:51, Eugen Block a écrit :
In Slack it was confirmed for 17.2.8, other versions have not been mentioned (yet).
Zitat von "Fox, Kevin M" <Kevin.Fox@pnnl.gov>:
One of the referenced issues seems to indicate it may be in quincy too? does anyone know if thats true, and when it was introduced? Does it go back even farther?
Thanks, Kevin
________________________________________ From: Devender Singh <devender@netskrt.io> Sent: Thursday, May 1, 2025 11:40 AM To: Alex Cc: Dan van der Ster; ceph-users Subject: [ceph-users] Re: Heads up -- 18.2.6 regression
Check twice before you click! This email originated from outside PNNL.
Hello Alex
As per Dan bug is in Reef which is v18.2.6 in open ceph.
I also upgraded my cluster to 18.2.6 before I saw the first message of this mail chain and I am having 120osds but yet I have not seen any issue where’s as one of my host remain down for 24hrs with 24 osds on it and now I joined back and all osds came active.
Regards Dev
On Thu, 1 May 2025 at 9:51 AM, Alex <mr.alexey@gmail.com> wrote:
Thanks.
According to Red Hat
Ceph 6 is Quincy Ceph 7 is Reef Ceph 8 is Squid
Is the bug in Reef or Squid? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi all, Just saw this thread - so a quick update on our experience here. We've been testing Quincy 17.2.8 and a couple of weeks ago we ran into this issue (before I heard about it anywhere on the mailing list). Our use case was cephfs on relatively full OSDs backed by NVMe. During fairly heavily parallel tests (up to 24 client nodes, 16 way parallel per node) a few dozen OSDs crashed with this BlueFS error. We were able to reproduce the problem at will. We didn't see any data corruption - once the OSDs got restarted (after stopping the test), they recovered without any side effects we could see. I ran into the tracker and the corresponding PR googling the error message. After I applied the patch in the referenced PR on top of 17.2.8 about a week ago (and rebuilt ceph-osd from source), and a solid week of testing has not produced any BlueFS issues on the same cluster - so that gives me high confidence that the patch resolves the issue. This is all Quincy - we haven't tried Reef yet (this was one of a few surprises Quincy handed us on our path to upgrade from Pacific). Andras On 4/30/25 1:30 PM, Dan van der Ster wrote:
Hi all,
Just a quick heads up -- 18.2.6 has a bluefs regression [1] which was not present in 18.2.4. Please avoid upgrading until further notice.
Regards, Dan
[1] Introduced in https://tracker.ceph.com/issues/65356 and fixed in https://tracker.ceph.com/issues/69764
participants (9)
-
Alex
-
Andras Pataki
-
Dan van der Ster
-
Devender Singh
-
Enrico Bocchi
-
Eugen Block
-
Fox, Kevin M
-
Frédéric Nass
-
Michel Jouvin