Hi Reed,
Thankyou so much  for the input and support. We have tried using the variable suggested by you, but could not see any impact on the current system. 
"ceph fs set cephfs allow_standby_replay true " it did not create any impact in the failover time

Furthermore we have tried more scenarios that we tested using our test :
scenario 1:
image.png

In both the test cases above we saw some extra delay of around 15 seconds + 8-10 seconds. (total 21-25 seconds for failover in case of power-off/reboot), 

Query: Any specific config that may need to be tweaked/tried to reduce this time for MDS to know that it has to activate and start the standby MDS Node?)

  Scenario 2:  

Please suggest/advice if we can try to configure to achieve minimal failover duration in the first two scenarios. 

Best Regards,
Lokendra




On Thu, Apr 29, 2021 at 1:47 AM Reed Dier <reed.dier@focusvq.com> wrote:
I don't have anything of merit to add to this, but it would be an interesting addition to your testing to see if active+standby-replay makes any difference with test-case1.

I don't think it would be applicable to any of the other use-cases, as a standby-replay MDS is bound to a single rank, meaning its bound to a single active MDS, and can't function as a standby for active:active.

https://docs.ceph.com/en/latest/cephfs/standby/#configuring-standby-replay

https://access.redhat.com/documentation/en-us/red_hat_ceph_storage/2/html/ceph_file_system_guide_technology_preview/installing_and_configuring_ceph_metadata_servers_mds#mds-configuring-standby-daemons-standby-replay

Good luck and look forward to hearing feedback/more results.

Reed

On Apr 27, 2021, at 8:40 AM, Lokendra Rathour <lokendrarathour@gmail.com> wrote:

Hi Team,
We have setup two Node Ceph Cluster using Native Cephfs Driver with Details as:
  • 3 Node / 2 Node MDS Cluster
  • 3 Node Monitor Quorum
  • 2 Node OSD
  • 2 Nodes for Manager

 

Cephnode3 have only Mon and MDS (only for test case 4-7) rest two nodes i.e. cephnode1 and cephnode2 have (mgr,mds,mon,rgw)

 

We have tested following failover scenarios for Native Cephfs Driver by mounting for any one sub-volume on a VM or client with continuous I/O operations(Directory creation after every 1 Second):

<image.png>


In the table above we have few queries as:
  • Refer test case 2 and test case 7, both are similar test case with only difference in number of Ceph MDS with time for both the test cases is different. It should be zero. But time is coming as 17 seconds for testcase 7.
  • Is there any configurable parameter/any configuration which we need to make in the Ceph cluster to get the failover time reduced to few seconds?
In current default deployment we are getting something around 35-40 seconds.

 

 

 

Best Regards,

--
~ Lokendra
skype: lokendrarathour


_______________________________________________
ceph-users mailing list -- ceph-users@ceph.io
To unsubscribe send an email to ceph-users-leave@ceph.io



--
~ Lokendra
www.inertiaspeaks.com
www.inertiagroups.com
skype: lokendrarathour