I haven't gone through this thread yet but I want to note for those reading that we do now have documentation (thanks for the frequent pokes Janek!) for the recall configurations:
https://docs.ceph.com/en/latest/cephfs/cache-configuration/#mds-recall
Please let us know if it's missing information or if something could be more clear. The documentation has helped a big deal already and I've been playing around with them quite a bit recently. What's missing, obviously, are recommended settings for individual scenarios (at least ballparks). But
Hi Patrick, that is hard to come by without experimenting first (I wouldn't call our deployment massive, but very likely significantly above average and I don't know what scale the developers are usually testing at). As I mentioned in the other thread, I am testing Dan's recommendations at the moment and will refine them for our purposes. The effects of individual tweaks are hard to assess without dedicated benchmarks (although "MDS not hanging up" is already somewhat of a benchmark :-)).