Hi, I'm not sure if it could be related to summer vacation, but I feel like there should have been several reports similar to yours by now. I don't have upgraded any production cluster to 20 yet, so I can't really deny or confirm anything. Did you upgrade the OS as well? Maybe there were some relevant changes there. Or some non-persistent configs were overwritten, either in the OS or Ceph. Was there more deep-scrubbing after the upgrade which could have impacted the IO? Regards, Eugen Zitat von Moti via ceph-users <ceph-users@ceph.io>:
Hello everyone! I recently upgraded from 18.2.4 to 20.2.1 (so it might be a Squid thing as well as Tentacle) and latency went ~3-4x more on PUT, even worse on GET and HEAD requests.
There is no resource cap or significant increase in usage is visible. Neither fast ec nor any other new features are enabled yet, just did the upgrade and left it as is. For measuring latency I have a services that probes the cluster with consistent load so it's not a difference in load either. PG status history does not show any pg operation being heavily in process.
Cluster has a hybrid build so OSDs have HDD for data and NVMe for db block. Both ec and replicated pools exist. The strange hint was: NVMe disks' iops went from something around 5-6K constant io per second to almost zero (less than a hundred per second, on all host/OSDs, with some non-frequent high peaks), while the HDD iops is not much changed. db blocks on NVMe are still attached to OSDs with no issue.
I can't find a clue to explain this. Got the same results on 3 separate clusters (all went version 18 --> 20, same hybrid build and configs). Nothing in config dump changes (changed defaults) catches my attention to have caused this.
Anyone had similar issues? Any tips and ideas on what to check? _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io