Have a problem with haproxy/keepalived/ganesha/docker
Hello! I've installed my 5-node CEPH cluster next install NFS server by command: ceph nfs cluster create nfshacluster 5 --ingress --virtual_ip 192.168.171.48/26 --ingress-mode haproxy-protocol. I don't understand fully how this must be works but when i stop NFS daemon even on one of this nodes I've see that writing on NFS shares is disappear (testing via vdbench). As i understand it is wrong and IO from stopped daemon must switching to another NFS daemon without any impact on IO. Can someone help me with troubleshoot this issue? Or explain how done full-fledged Active-Active HA NFS Cluster for production use. Thanks! Руслан Нурабаев Старший инженер Сектор ИТ платформы Отдел развития опорной сети Департамент развития сети +77012119272 Ruslan.Nurabayev@kcell.kz -----Original Message----- From: ceph-users-request@ceph.io <ceph-users-request@ceph.io> Sent: Thursday, April 11, 2024 15:07 To: Ruslan Nurabayev <Ruslan.Nurabayev@kcell.kz> Subject: Welcome to the "ceph-users" mailing list [You don't often get email from ceph-users-request@ceph.io. Learn why this is important at https://aka.ms/LearnAboutSenderIdentification ] Welcome to the "ceph-users" mailing list! To post to this list, send your email to: ceph-users@ceph.io You can unsubscribe or make adjustments to your options via email by sending a message to: ceph-users-request@ceph.io with the word 'help' in the subject or body (don't include the quotes), and you will get back a message with instructions. You will need your password to change your options, but for security purposes, this password is not included here. If you have forgotten your password you will need to reset it via the web UI. ________________________________ **************************************************************************************** Осы хабарлама және онымен берілетін кез келген файлдар құпия болып табылады және олар мекенжайда көрсетілген жеке немесе заңды тұлғалардың пайдалануына ғана арналған. Егер сіз болжамды алушы болып табылмайтын болсаңыз, осы арқылы осындай ақпаратты кез келген таратуға, жіберуге, көшіруге немесе пайдалануға қатаң тыйым салынатыны және осы электрондық хабарлама дереу жойылуға тиіс екендігін хабарлаймыз. KCELL осы хабарламадағы кез келген ақпараттың дәлдігіне немесе толықтығына қатысты ешқандай кепілдік бермейді және сол арқылы онда қамтылған ақпарат үшін немесе оны беру, қабылдау, сақтау немесе қандай да бір түрде пайдалану үшін кез келген жауапкершілікті болдырмайды. Осы хабарламада айтылған пікірлер тек жіберушіге ғана тиесілі және KCELL пікірін де білдіруі міндетті емес. Бұл электрондық хабарлама барлық танымал компьютерлік вирустарға тексерілді. **************************************************************************************** Данное сообщение и любые передаваемые с ним файлы являются конфиденциальными и предназначены исключительно для использования физическими или юридическими лицами, которым они адресованы. Если вы не являетесь предполагаемым получателем, настоящим уведомляем о том, что любое распространение, пересылка, копирование или использование такой информации строго запрещено, и данное электронное сообщение должно быть немедленно удалено. KCELL не дает никаких гарантий относительно точности или полноты любой информации, содержащейся в данном сообщении, и тем самым исключает любую ответственность за информацию, содержащуюся в нем, или за ее передачу, прием, хранение или использование каким-либо образом. Мнения, выраженные в данном сообщении, принадлежат только отправителю и не обязательно отражают мнение KCELL. Данное электронное сообщение было проверено на наличие всех известных компьютерных вирусов. **************************************************************************************** This e-mail and any files transmitted with it are confidential and intended solely for the use of the individual or entity to whom they are addressed. If you are not the intended recipient you are hereby notified that any dissemination, forwarding, copying or use of any of the information is strictly prohibited, and the e-mail should immediately be deleted. KCELL makes no warranty as to the accuracy or completeness of any information contained in this message and hereby excludes any liability of any kind for the information contained therein or for the information transmission, reception, storage or use of such in any way whatsoever. The opinions expressed in this message belong to sender alone and may not necessarily reflect the opinions of KCELL. This e-mail has been scanned for all known computer viruses. ****************************************************************************************
Hi, I believe I can confirm your suspicion, I have a test cluster on Reef 18.2.1 and deployed nfs without HAProxy but with keepalived [1]. Stopping the active NFS daemon doesn't trigger anything, the MGR notices that it's stopped at some point, but nothing else seems to happen. I didn't wait too long, but I suspect that it's not supposed to work like that. I'll check tracker for existing issues. Thanks, Eugen [1] https://docs.ceph.com/en/reef/cephadm/services/nfs/#nfs-with-virtual-ip-but-... Zitat von Ruslan Nurabayev <Ruslan.Nurabayev@kcell.kz>:
Hello! I've installed my 5-node CEPH cluster next install NFS server by command: ceph nfs cluster create nfshacluster 5 --ingress --virtual_ip 192.168.171.48/26 --ingress-mode haproxy-protocol. I don't understand fully how this must be works but when i stop NFS daemon even on one of this nodes I've see that writing on NFS shares is disappear (testing via vdbench). As i understand it is wrong and IO from stopped daemon must switching to another NFS daemon without any impact on IO. Can someone help me with troubleshoot this issue? Or explain how done full-fledged Active-Active HA NFS Cluster for production use.
Thanks!
Руслан Нурабаев Старший инженер Сектор ИТ платформы Отдел развития опорной сети Департамент развития сети
+77012119272 Ruslan.Nurabayev@kcell.kz
-----Original Message----- From: ceph-users-request@ceph.io <ceph-users-request@ceph.io> Sent: Thursday, April 11, 2024 15:07 To: Ruslan Nurabayev <Ruslan.Nurabayev@kcell.kz> Subject: Welcome to the "ceph-users" mailing list
[You don't often get email from ceph-users-request@ceph.io. Learn why this is important at https://aka.ms/LearnAboutSenderIdentification ]
Welcome to the "ceph-users" mailing list!
To post to this list, send your email to:
ceph-users@ceph.io
You can unsubscribe or make adjustments to your options via email by sending a message to:
ceph-users-request@ceph.io
with the word 'help' in the subject or body (don't include the quotes), and you will get back a message with instructions. You will need your password to change your options, but for security purposes, this password is not included here. If you have forgotten your password you will need to reset it via the web UI.
________________________________
**************************************************************************************** Осы хабарлама және онымен берілетін кез келген файлдар құпия болып табылады және олар мекенжайда көрсетілген жеке немесе заңды тұлғалардың пайдалануына ғана арналған. Егер сіз болжамды алушы болып табылмайтын болсаңыз, осы арқылы осындай ақпаратты кез келген таратуға, жіберуге, көшіруге немесе пайдалануға қатаң тыйым салынатыны және осы электрондық хабарлама дереу жойылуға тиіс екендігін хабарлаймыз. KCELL осы хабарламадағы кез келген ақпараттың дәлдігіне немесе толықтығына қатысты ешқандай кепілдік бермейді және сол арқылы онда қамтылған ақпарат үшін немесе оны беру, қабылдау, сақтау немесе қандай да бір түрде пайдалану үшін кез келген жауапкершілікті болдырмайды. Осы хабарламада айтылған пікірлер тек жіберушіге ғана тиесілі және KCELL пікірін де білдіруі міндетті емес. Бұл электрондық хабарлама барлық танымал компьютерлік вирустарға тексерілді. **************************************************************************************** Данное сообщение и любые передаваемые с ним файлы являются конфиденциальными и предназначены исключительно для использования физическими или юридическими лицами, которым они адресованы. Если вы не являетесь предполагаемым получателем, настоящим уведомляем о том, что любое распространение, пересылка, копирование или использование такой информации строго запрещено, и данное электронное сообщение должно быть немедленно удалено. KCELL не дает никаких гарантий относительно точности или полноты любой информации, содержащейся в данном сообщении, и тем самым исключает любую ответственность за информацию, содержащуюся в нем, или за ее передачу, прием, хранение или использование каким-либо образом. Мнения, выраженные в данном сообщении, принадлежат только отправителю и не обязательно отражают мнение KCELL. Данное электронное сообщение было проверено на наличие всех известных компьютерных вирусов. **************************************************************************************** This e-mail and any files transmitted with it are confidential and intended solely for the use of the individual or entity to whom they are addressed. If you are not the intended recipient you are hereby notified that any dissemination, forwarding, copying or use of any of the information is strictly prohibited, and the e-mail should immediately be deleted. KCELL makes no warranty as to the accuracy or completeness of any information contained in this message and hereby excludes any liability of any kind for the information contained therein or for the information transmission, reception, storage or use of such in any way whatsoever. The opinions expressed in this message belong to sender alone and may not necessarily reflect the opinions of KCELL. This e-mail has been scanned for all known computer viruses. **************************************************************************************** _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi, On 16.04.24 10:49, Eugen Block wrote:
I believe I can confirm your suspicion, I have a test cluster on Reef 18.2.1 and deployed nfs without HAProxy but with keepalived [1]. Stopping the active NFS daemon doesn't trigger anything, the MGR notices that it's stopped at some point, but nothing else seems to happen.
There is currently no failover for NFS. The ingress service (haproxy + keepalived) that cephadm deploys for an NFS cluster does not have a health check configured. Haproxy does not notice if a backend NFS server dies. This does not matter as there is no failover and the NFS client cannot be "load balanced" to another backend NFS server. There is no use to configure an ingress service currently without failover. The NFS clients have to remount the NFS share in case of their current NFS server dies anyway. Regards -- Robert Sander Heinlein Consulting GmbH Schwedter Str. 8/9b, 10119 Berlin http://www.heinlein-support.de Tel: 030 / 405051-43 Fax: 030 / 405051-19 Zwangsangaben lt. §35a GmbHG: HRB 220009 B / Amtsgericht Berlin-Charlottenburg, Geschäftsführer: Peer Heinlein -- Sitz: Berlin
Ah, okay, thanks for the hint. In that case what I see is expected. Zitat von Robert Sander <r.sander@heinlein-support.de>:
Hi,
On 16.04.24 10:49, Eugen Block wrote:
I believe I can confirm your suspicion, I have a test cluster on Reef 18.2.1 and deployed nfs without HAProxy but with keepalived [1]. Stopping the active NFS daemon doesn't trigger anything, the MGR notices that it's stopped at some point, but nothing else seems to happen.
There is currently no failover for NFS.
The ingress service (haproxy + keepalived) that cephadm deploys for an NFS cluster does not have a health check configured. Haproxy does not notice if a backend NFS server dies. This does not matter as there is no failover and the NFS client cannot be "load balanced" to another backend NFS server.
There is no use to configure an ingress service currently without failover.
The NFS clients have to remount the NFS share in case of their current NFS server dies anyway.
Regards -- Robert Sander Heinlein Consulting GmbH Schwedter Str. 8/9b, 10119 Berlin
http://www.heinlein-support.de
Tel: 030 / 405051-43 Fax: 030 / 405051-19
Zwangsangaben lt. §35a GmbHG: HRB 220009 B / Amtsgericht Berlin-Charlottenburg, Geschäftsführer: Peer Heinlein -- Sitz: Berlin _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Question about HA here: I understood the documentation of the fuse NFS client such that the connection state of all NFS clients is stored on ceph in rados objects and, if using a floating IP, the NFS clients should just recover from a short network timeout. Not sure if this is what should happen with this specific HA set-up in the original request, but a fail-over of the NFS server ought to be handled gracefully by starting a new one up with the IP of the down one. Or not? Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14 ________________________________________ From: Eugen Block <eblock@nde.ag> Sent: Tuesday, April 16, 2024 11:24 AM To: ceph-users@ceph.io Subject: [ceph-users] Re: Have a problem with haproxy/keepalived/ganesha/docker Ah, okay, thanks for the hint. In that case what I see is expected. Zitat von Robert Sander <r.sander@heinlein-support.de>:
Hi,
On 16.04.24 10:49, Eugen Block wrote:
I believe I can confirm your suspicion, I have a test cluster on Reef 18.2.1 and deployed nfs without HAProxy but with keepalived [1]. Stopping the active NFS daemon doesn't trigger anything, the MGR notices that it's stopped at some point, but nothing else seems to happen.
There is currently no failover for NFS.
The ingress service (haproxy + keepalived) that cephadm deploys for an NFS cluster does not have a health check configured. Haproxy does not notice if a backend NFS server dies. This does not matter as there is no failover and the NFS client cannot be "load balanced" to another backend NFS server.
There is no use to configure an ingress service currently without failover.
The NFS clients have to remount the NFS share in case of their current NFS server dies anyway.
Regards -- Robert Sander Heinlein Consulting GmbH Schwedter Str. 8/9b, 10119 Berlin
http://www.heinlein-support.de
Tel: 030 / 405051-43 Fax: 030 / 405051-19
Zwangsangaben lt. §35a GmbHG: HRB 220009 B / Amtsgericht Berlin-Charlottenburg, Geschäftsführer: Peer Heinlein -- Sitz: Berlin _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hm, no, I can't confirm it yet. I missed something in the config, the failover happens and a new nfs daemon is deployed on a different node. But I still see client interruptions so I'm gonna look into that first. Zitat von Eugen Block <eblock@nde.ag>:
Hi,
I believe I can confirm your suspicion, I have a test cluster on Reef 18.2.1 and deployed nfs without HAProxy but with keepalived [1]. Stopping the active NFS daemon doesn't trigger anything, the MGR notices that it's stopped at some point, but nothing else seems to happen. I didn't wait too long, but I suspect that it's not supposed to work like that. I'll check tracker for existing issues.
Thanks, Eugen
[1] https://docs.ceph.com/en/reef/cephadm/services/nfs/#nfs-with-virtual-ip-but-...
Zitat von Ruslan Nurabayev <Ruslan.Nurabayev@kcell.kz>:
Hello! I've installed my 5-node CEPH cluster next install NFS server by command: ceph nfs cluster create nfshacluster 5 --ingress --virtual_ip 192.168.171.48/26 --ingress-mode haproxy-protocol. I don't understand fully how this must be works but when i stop NFS daemon even on one of this nodes I've see that writing on NFS shares is disappear (testing via vdbench). As i understand it is wrong and IO from stopped daemon must switching to another NFS daemon without any impact on IO. Can someone help me with troubleshoot this issue? Or explain how done full-fledged Active-Active HA NFS Cluster for production use.
Thanks!
Руслан Нурабаев Старший инженер Сектор ИТ платформы Отдел развития опорной сети Департамент развития сети
+77012119272 Ruslan.Nurabayev@kcell.kz
-----Original Message----- From: ceph-users-request@ceph.io <ceph-users-request@ceph.io> Sent: Thursday, April 11, 2024 15:07 To: Ruslan Nurabayev <Ruslan.Nurabayev@kcell.kz> Subject: Welcome to the "ceph-users" mailing list
[You don't often get email from ceph-users-request@ceph.io. Learn why this is important at https://aka.ms/LearnAboutSenderIdentification ]
Welcome to the "ceph-users" mailing list!
To post to this list, send your email to:
ceph-users@ceph.io
You can unsubscribe or make adjustments to your options via email by sending a message to:
ceph-users-request@ceph.io
with the word 'help' in the subject or body (don't include the quotes), and you will get back a message with instructions. You will need your password to change your options, but for security purposes, this password is not included here. If you have forgotten your password you will need to reset it via the web UI.
________________________________
**************************************************************************************** Осы хабарлама және онымен берілетін кез келген файлдар құпия болып табылады және олар мекенжайда көрсетілген жеке немесе заңды тұлғалардың пайдалануына ғана арналған. Егер сіз болжамды алушы болып табылмайтын болсаңыз, осы арқылы осындай ақпаратты кез келген таратуға, жіберуге, көшіруге немесе пайдалануға қатаң тыйым салынатыны және осы электрондық хабарлама дереу жойылуға тиіс екендігін хабарлаймыз. KCELL осы хабарламадағы кез келген ақпараттың дәлдігіне немесе толықтығына қатысты ешқандай кепілдік бермейді және сол арқылы онда қамтылған ақпарат үшін немесе оны беру, қабылдау, сақтау немесе қандай да бір түрде пайдалану үшін кез келген жауапкершілікті болдырмайды. Осы хабарламада айтылған пікірлер тек жіберушіге ғана тиесілі және KCELL пікірін де білдіруі міндетті емес. Бұл электрондық хабарлама барлық танымал компьютерлік вирустарға тексерілді. **************************************************************************************** Данное сообщение и любые передаваемые с ним файлы являются конфиденциальными и предназначены исключительно для использования физическими или юридическими лицами, которым они адресованы. Если вы не являетесь предполагаемым получателем, настоящим уведомляем о том, что любое распространение, пересылка, копирование или использование такой информации строго запрещено, и данное электронное сообщение должно быть немедленно удалено. KCELL не дает никаких гарантий относительно точности или полноты любой информации, содержащейся в данном сообщении, и тем самым исключает любую ответственность за информацию, содержащуюся в нем, или за ее передачу, прием, хранение или использование каким-либо образом. Мнения, выраженные в данном сообщении, принадлежат только отправителю и не обязательно отражают мнение KCELL. Данное электронное сообщение было проверено на наличие всех известных компьютерных вирусов. **************************************************************************************** This e-mail and any files transmitted with it are confidential and intended solely for the use of the individual or entity to whom they are addressed. If you are not the intended recipient you are hereby notified that any dissemination, forwarding, copying or use of any of the information is strictly prohibited, and the e-mail should immediately be deleted. KCELL makes no warranty as to the accuracy or completeness of any information contained in this message and hereby excludes any liability of any kind for the information contained therein or for the information transmission, reception, storage or use of such in any way whatsoever. The opinions expressed in this message belong to sender alone and may not necessarily reflect the opinions of KCELL. This e-mail has been scanned for all known computer viruses. **************************************************************************************** _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I used docker installation with default configuration. And I googled my problem and saw at reddit this post: https://www.reddit.com/r/ceph/comments/10bcwra/nfs_cluster_ha_not_working_wh... I added this configuration with "check" option in my installation. Now haproxy detects fail of node but vdbench show writes is lost even if one node from 5 is down. Also I used mount -o vers=4.1 which is more stable than 4.2 as I see. How can i set up true HA before using this in a production?
participants (6)
-
Eugen Block
-
Frank Schilder
-
Marc
-
Robert Sander
-
Rusik NV
-
Ruslan Nurabayev