Hi Frank, On 5/11/20 3:03 PM, Frank Schilder wrote:
OK, the command finally executed and it looks like the cluster is running stable for now. However, I'm afraid that 90s might not be sustainable.
Questions: Can I leave the beacon_grace at 90s? Is there a better parameter to set? Why is the MGR getting overloaded on a rather small cluster with 160 OSDs? How does this scale?
I wonder if https://tracker.ceph.com/issues/45439 might be related to what you're observing here? In this issue, Andras suggests: "Increasing mgr_stats_period to 15 seconds reduces the load and brings ceph-mgr back to responsive again." Maybe that helps? Lenz -- SUSE Software Solutions Germany GmbH - Maxfeldstr. 5 - 90409 Nuernberg GF: Felix Imendörffer, HRB 36809 (AG Nürnberg)