Identify laggy PGs
Hi, currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones. Any tooling to identify or count PG performance and map that to OSDs? -- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal.
We either facing that. Have a look in the logs reported failed osd. I count the occurrence and offline compact those, it can help for a while. Normally for us compacting blocking the operation on it. Istvan ________________________________ From: Boris <bb@kervyn.de> Sent: Saturday, August 10, 2024 5:30:54 PM To: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [ceph-users] Identify laggy PGs Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________ Hi, currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones. Any tooling to identify or count PG performance and map that to OSDs? -- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io ________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses.
hmm.. will try that. Thanks Am Sa., 10. Aug. 2024 um 13:33 Uhr schrieb Szabo, Istvan (Agoda) < Istvan.Szabo@agoda.com>:
We either facing that. Have a look in the logs reported failed osd. I count the occurrence and offline compact those, it can help for a while. Normally for us compacting blocking the operation on it.
Istvan ------------------------------ *From:* Boris <bb@kervyn.de> *Sent:* Saturday, August 10, 2024 5:30:54 PM *To:* ceph-users@ceph.io <ceph-users@ceph.io> *Subject:* [ceph-users] Identify laggy PGs
Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________
Hi,
currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones.
Any tooling to identify or count PG performance and map that to OSDs?
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
------------------------------ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses.
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal.
Hi, how big are those PGs? If they're huge and are deep-scrubbed, for example, that can cause significant delays. I usually look at 'ceph pg ls-by-pool {pool}' and the "BYTES" column. Zitat von Boris <bb@kervyn.de>:
Hi,
currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones.
Any tooling to identify or count PG performance and map that to OSDs?
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
PGs are roughtly 35GB. Am Mi., 14. Aug. 2024 um 09:25 Uhr schrieb Eugen Block <eblock@nde.ag>:
Hi,
how big are those PGs? If they're huge and are deep-scrubbed, for example, that can cause significant delays. I usually look at 'ceph pg ls-by-pool {pool}' and the "BYTES" column.
Zitat von Boris <bb@kervyn.de>:
Hi,
currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones.
Any tooling to identify or count PG performance and map that to OSDs?
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal.
Just curiously I've checked my pg size which is like 150GB, when are we talking about big pgs? ________________________________ From: Eugen Block <eblock@nde.ag> Sent: Wednesday, August 14, 2024 2:23 PM To: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [ceph-users] Re: Identify laggy PGs Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________ Hi, how big are those PGs? If they're huge and are deep-scrubbed, for example, that can cause significant delays. I usually look at 'ceph pg ls-by-pool {pool}' and the "BYTES" column. Zitat von Boris <bb@kervyn.de>:
Hi,
currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones.
Any tooling to identify or count PG performance and map that to OSDs?
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io ________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses.
Hi Boris,
PGs are roughtly 35GB.
that's not huge. You wrote you drained one OSD which helped with the flapping, so you don't have flapping OSDs anymore at all? If you have identified problematic PGs, you can get the OSD mapping like this: ceph pg map 26.7 osdmap e14121 pg 26.7 (26.7) -> up [2,5,8] acting [2,5,8]
Just curiously I've checked my pg size which is like 150GB, when are we talking about big pgs?
That depends on a couple of factors. Just one example from a customer cluster: They had 240 HDD OSDs with roughly 1 PB a couple of months ago, which resulted in PG sizes around 400 GB. This led to very long deep-scrubs, utilizing the OSDs for quite some time, with a noticable impact on their application. This was not only a performance issue but also a balancing issue: if the OSDs deviate by only 5 PGs that makes a difference of 2 TB, on 8 TB disks that is 25% which can quickly bring the OSDs to or above 85% usage. That's why we quadrupled the PGs for the main pool which improved balancing a lot, and deep-scrubs per PG also run much faster with the cost of having more PGs to scrub, of course. Zitat von "Szabo, Istvan (Agoda)" <Istvan.Szabo@agoda.com>:
Just curiously I've checked my pg size which is like 150GB, when are we talking about big pgs? ________________________________ From: Eugen Block <eblock@nde.ag> Sent: Wednesday, August 14, 2024 2:23 PM To: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [ceph-users] Re: Identify laggy PGs
Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________
Hi,
how big are those PGs? If they're huge and are deep-scrubbed, for example, that can cause significant delays. I usually look at 'ceph pg ls-by-pool {pool}' and the "BYTES" column.
Zitat von Boris <bb@kervyn.de>:
Hi,
currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones.
Any tooling to identify or count PG performance and map that to OSDs?
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses.
The current ceph recommendation is to use between 100-200 PGs/OSD. Therefore, a large PG is a PG that has more data than 0.5-1% of the disk capacity and you should split PGs for the relevant pool. A huge PG is a PG for which deep-scrub takes much longer than 20min on HDD and 4-5min on SSD. Average deep-scrub times (time it takes to deep-scrub) are actually a very good way of judging if PGs are too large. These times roughly correlate with the time it takes to copy a PG. On SSDs we aim for 200+PGs/OSD and for HDDs for 150PGs/OSD. For very large HDD disks (>=16TB) we consider raising this to 300PGs/OSD due to excessively long deep-scrub times per PG. Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14 ________________________________________ From: Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com> Sent: Wednesday, August 14, 2024 12:00 PM To: Eugen Block; ceph-users@ceph.io Subject: [ceph-users] Re: Identify laggy PGs Just curiously I've checked my pg size which is like 150GB, when are we talking about big pgs? ________________________________ From: Eugen Block <eblock@nde.ag> Sent: Wednesday, August 14, 2024 2:23 PM To: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [ceph-users] Re: Identify laggy PGs Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________ Hi, how big are those PGs? If they're huge and are deep-scrubbed, for example, that can cause significant delays. I usually look at 'ceph pg ls-by-pool {pool}' and the "BYTES" column. Zitat von Boris <bb@kervyn.de>:
Hi,
currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones.
Any tooling to identify or count PG performance and map that to OSDs?
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io ________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I just checked the logs, and there are also laggy PGs when all others are active+clean. After adding 5x8tb disks the laggy PGs stopped. ¯\_(ツ)_/¯ Maybe upping the PGs could do something do something to make deep scrubbing faster. But we have 2,4 and 8 TB disks in this cluster. And they are only sata SSDs, so not particularly fast ones. I always thought that too many PGs have impact on the disk IO. I guess this is wrong? So I could double the PGs in the pool and see if things become better. And yes, removing that single OSD from the cluster stopped the flapping of "monitor marked osd.N down".
Am 15.08.2024 um 10:14 schrieb Frank Schilder <frans@dtu.dk>:
The current ceph recommendation is to use between 100-200 PGs/OSD. Therefore, a large PG is a PG that has more data than 0.5-1% of the disk capacity and you should split PGs for the relevant pool.
A huge PG is a PG for which deep-scrub takes much longer than 20min on HDD and 4-5min on SSD.
Average deep-scrub times (time it takes to deep-scrub) are actually a very good way of judging if PGs are too large. These times roughly correlate with the time it takes to copy a PG.
On SSDs we aim for 200+PGs/OSD and for HDDs for 150PGs/OSD. For very large HDD disks (>=16TB) we consider raising this to 300PGs/OSD due to excessively long deep-scrub times per PG.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com> Sent: Wednesday, August 14, 2024 12:00 PM To: Eugen Block; ceph-users@ceph.io Subject: [ceph-users] Re: Identify laggy PGs
Just curiously I've checked my pg size which is like 150GB, when are we talking about big pgs? ________________________________ From: Eugen Block <eblock@nde.ag> Sent: Wednesday, August 14, 2024 2:23 PM To: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [ceph-users] Re: Identify laggy PGs
Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________
Hi,
how big are those PGs? If they're huge and are deep-scrubbed, for example, that can cause significant delays. I usually look at 'ceph pg ls-by-pool {pool}' and the "BYTES" column.
Zitat von Boris <bb@kervyn.de>:
Hi,
currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones.
Any tooling to identify or count PG performance and map that to OSDs?
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I always thought that too many PGs have impact on the disk IO. I guess this is wrong?
Mostly when they’re spinners. Especially back in the Filestore days with a colocated journal. Don’t get me started on that. Too many PGs can exhaust RAM if you’re tight - or using Filestore still. For a SATA SSD I’d set pg_nums to average 200-300 per drive. Your size mix complicates, though, because the larger OSDs will get many more than the smaller. Be sure to set mon_max_pg_per_osd to like 1000. You might be experiment with primary affinity, so that the smaller OSDs are more likely to be primaries and thus will get more load. I’ve seen a first-order approximation here increase read throughput by 20% To
So I could double the PGs in the pool and see if things become better.
And yes, removing that single OSD from the cluster stopped the flapping of "monitor marked osd.N down".
Am 15.08.2024 um 10:14 schrieb Frank Schilder <frans@dtu.dk>:
The current ceph recommendation is to use between 100-200 PGs/OSD. Therefore, a large PG is a PG that has more data than 0.5-1% of the disk capacity and you should split PGs for the relevant pool.
A huge PG is a PG for which deep-scrub takes much longer than 20min on HDD and 4-5min on SSD.
Average deep-scrub times (time it takes to deep-scrub) are actually a very good way of judging if PGs are too large. These times roughly correlate with the time it takes to copy a PG.
On SSDs we aim for 200+PGs/OSD and for HDDs for 150PGs/OSD. For very large HDD disks (>=16TB) we consider raising this to 300PGs/OSD due to excessively long deep-scrub times per PG.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Szabo, Istvan (Agoda) <Istvan.Szabo@agoda.com> Sent: Wednesday, August 14, 2024 12:00 PM To: Eugen Block; ceph-users@ceph.io Subject: [ceph-users] Re: Identify laggy PGs
Just curiously I've checked my pg size which is like 150GB, when are we talking about big pgs? ________________________________ From: Eugen Block <eblock@nde.ag> Sent: Wednesday, August 14, 2024 2:23 PM To: ceph-users@ceph.io <ceph-users@ceph.io> Subject: [ceph-users] Re: Identify laggy PGs
Email received from the internet. If in doubt, don't click any link nor open any attachment ! ________________________________
Hi,
how big are those PGs? If they're huge and are deep-scrubbed, for example, that can cause significant delays. I usually look at 'ceph pg ls-by-pool {pool}' and the "BYTES" column.
Zitat von Boris <bb@kervyn.de>:
Hi,
currently we encouter laggy PGs and I would like to find out what is causing it. I suspect it might be one or more failing OSDs. We had flapping OSDs and I synced one out, which helped with the flapping, but it doesn't help with the laggy ones.
Any tooling to identify or count PG performance and map that to OSDs?
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
________________________________ This message is confidential and is for the sole use of the intended recipient(s). It may also be privileged or otherwise protected by copyright or other legal rules. If you have received it by mistake please let us know by reply email and delete it from your system. It is prohibited to copy this message or disclose its content to anyone. Any confidentiality or privilege is not waived or lost by any mistaken delivery or unauthorized disclosure of the message. All messages sent to and from Agoda may be monitored to ensure compliance with company policies, to protect the company's interests and to remove potential malware. Electronic messages may be intercepted, amended, lost or deleted, or contain viruses. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Good to know. Everything is bluestore and usually 5 spinners share an SSD for block.db. Memory should not be a problem. We plan with 4GB / OSD with a minimum of 256GB memory. The pirmary affinity is a nice idea. I only thought about it in our s3 cluster, because the index is on SAS AND SATA SSDs and I use the SAS as primary and the sata only for replication. Am Sa., 17. Aug. 2024 um 15:23 Uhr schrieb Anthony D'Atri < aad@dreamsnake.net>:
Mostly when they’re spinners. Especially back in the Filestore days with a colocated journal. Don’t get me started on that.
Too many PGs can exhaust RAM if you’re tight - or using Filestore still.
For a SATA SSD I’d set pg_nums to average 200-300 per drive. Your size mix complicates, though, because the larger OSDs will get many more than the smaller. Be sure to set mon_max_pg_per_osd to like 1000.
You might be experiment with primary affinity, so that the smaller OSDs are more likely to be primaries and thus will get more load. I’ve seen a first-order approximation here increase read throughput by 20%
-- Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im groüen Saal.
participants (5)
-
Anthony D'Atri
-
Boris
-
Eugen Block
-
Frank Schilder
-
Szabo, Istvan (Agoda)