monitor exclusive lock when rbd client died abruptly
Hi all, When the client writes to rbd image, it'll hold the "exclusive lock". If the client dies abruptly without releaseing the "exclusive lock", how the other client know that they could write to the rbd image? I write below program. If the first process/client died abrutply (built with -DKILL_DEAD, send kill signal), the second process/client(build without -DKILL_DEAD) could still write to the rbd image. However, there's an abvious "time delay" that the second client could write to the rbd image. Is there's a mechanism to watch/monitor that the "exclusive lock" is released after the client, which hold it, died abruptly? Progam: 1 #include <rbd/librbd.hpp> 2 #include <rados/librados.hpp> 3 4 #include <cstring> 5 #include <iostream> 6 #include <string> 7 8 void err_msg(int ret, const std::string &msg = "") { 9 std::cerr << "[error] msg:" << msg << " strerror: " 10 << strerror(-ret) << std::endl; 11 } 12 void err_exit(int ret, const std::string &msg = "") { 13 err_msg(ret, msg); 14 exit(EXIT_FAILURE); 15 } 16 17 int main(int argc, char* argv[]) { 18 int ret = 0; 19 librados::Rados rados; 20 21 ret = rados.init("admin"); 22 if (ret < 0) 23 err_exit(ret,"failed to initialize rados"); 24 ret = rados.conf_read_file("ceph.conf"); 25 if (ret < 0) 26 err_exit(ret, "failed to parse ceph.conf"); 27 28 ret = rados.connect(); 29 if (ret < 0) 30 err_exit(ret, "failed to connect to rados cluster"); 31 32 librados::IoCtx io_ctx; 33 std::string pool_name = "rbd"; 34 ret = rados.ioctx_create(pool_name.c_str(), io_ctx); 35 if (ret < 0) { 36 rados.shutdown(); 37 err_exit(ret, "failed to create ioctx"); 38 } 39 40 // rbd 41 librbd::RBD rbd; 42 43 librbd::Image image; 44 std::string image_name = "fio_test"; 45 ret = rbd.open(io_ctx, image, image_name.c_str()); 46 if (ret < 0) { 47 io_ctx.close(); 48 rados.shutdown(); 49 err_exit(ret, "failed to open rbd image"); 50 } else { 51 std::cout << "open image succeeded" << std::endl; 52 } 53 54 ceph::bufferlist bw; 55 bw.append(std::string("changcheng")); 56 image.write(0, bw.length(), bw); 57 58 ceph::bufferlist br; 59 int read = image.read(0, bw.length(), br); 60 br.append(std::string(4, '\0')); 61 std::cout << br.c_str() << std::endl; 62 63 #if defined(KILL_DEAD) 64 while(1); 65 #endif 66 67 done: 68 image.close(); 69 io_ctx.close(); 70 rados.shutdown(); 71 exit(EXIT_SUCCESS); 72 } B.R. Changcheng
On Sun, Jul 19, 2020 at 3:22 PM Liu, Changcheng <changcheng.liu@intel.com> wrote:
Hi all, When the client writes to rbd image, it'll hold the "exclusive lock". If the client dies abruptly without releaseing the "exclusive lock", how the other client know that they could write to the rbd image?
I write below program. If the first process/client died abrutply (built with -DKILL_DEAD, send kill signal), the second process/client(build without -DKILL_DEAD) could still write to the rbd image. However, there's an abvious "time delay" that the second client could write to the rbd image.
Is there's a mechanism to watch/monitor that the "exclusive lock" is released after the client, which hold it, died abruptly?
Hi Changcheng, No, this time delay is it. Exclusive lock is built on RADOS watch/notify framework and the client is considered alive as long as its watch is present. When you kill the client, it stops responding to watch pings and the OSD times out the watch after 30 seconds. At that point the other client can break the lock (after adding the original client to the OSD blacklist). Thanks, Ilya
Hi all, I've checked below document: https://docs.ceph.com/docs/master/rbd/rbd-exclusive-locks/ The content is contradictory to the result shown in below experiment: The experient shows that another process could still write data to rbd volume while there already has a process write to the same rbd volume continuously. 1. ceph master head commit: commit 1dd932a8f565dc74cf2441ab139aa173575c0e92 Date: Thu Jul 16 09:42:30 2020 +0900 2. setup cluster build$ OSD=3 MON=1 MGR=1 RGW=0 MDS=0 ../src/vstart.sh -k -d -n 3. create rbd volume build$ bin/ceph -c ceph.conf osd pool create rbd 128 build$ bin/ceph -c ceph.conf osd pool application enable rbd rbd build$ bin/rbd -c ceph.conf create fio_test --size 10G --image-format=2 --rbd_default_features=13 4. set environment variable build$ export LD_LIBRARY_PATH=/home/nstcc3/work/src/ceph/build/lib:$LD_LIBRARY_PATH 5. build attached file build$ g++ librbdtest.cpp -DKILL_DEAD -I../src/include -L lib/ -lrados -lrbd -o first build$ g++ librbdtest.cpp -I../src/include -L lib/ -lrados -lrbd -o second 6. run program 1) build$ ./first "first" process write to rbd volume continuously. 2) build$ ./second "second" process could still write to the same volume and exit normally. source code file: librbdtest.cpp 1 #include <rbd/librbd.hpp> 2 #include <rados/librados.hpp> 3 4 #include <cstring> 5 #include <iostream> 6 #include <string> 7 #include <unistd.h> 8 9 void err_msg(int ret, const std::string &msg = "") { 10 std::cerr << "[error] msg:" << msg << " strerror: " 11 << strerror(-ret) << std::endl; 12 } 13 void err_exit(int ret, const std::string &msg = "") { 14 err_msg(ret, msg); 15 exit(EXIT_FAILURE); 16 } 17 18 int main(int argc, char* argv[]) { 19 #if !defined(KILL_DEAD) 20 sleep(30); 21 #endif 22 int ret = 0; 23 librados::Rados rados; 24 25 ret = rados.init("admin"); 26 if (ret < 0) 27 err_exit(ret,"failed to initialize rados"); 28 ret = rados.conf_read_file("ceph.conf"); 29 if (ret < 0) 30 err_exit(ret, "failed to parse ceph.conf"); 31 32 ret = rados.connect(); 33 if (ret < 0) 34 err_exit(ret, "failed to connect to rados cluster"); 35 36 librados::IoCtx io_ctx; 37 std::string pool_name = "rbd"; 38 ret = rados.ioctx_create(pool_name.c_str(), io_ctx); 39 if (ret < 0) { 40 rados.shutdown(); 41 err_exit(ret, "failed to create ioctx"); 42 } 43 44 // rbd 45 librbd::RBD rbd; 46 47 librbd::Image image; 48 std::string image_name = "fio_test"; 49 ret = rbd.open(io_ctx, image, image_name.c_str()); 50 if (ret < 0) { 51 io_ctx.close(); 52 rados.shutdown(); 53 err_exit(ret, "failed to open rbd image"); 54 } else { 55 std::cout << "open image succeeded" << std::endl; 56 } 57 58 ceph::bufferlist bw; 59 #if !defined(KILL_DEAD) 60 bw.append(std::string("checkcheng")); 61 #else 62 bw.append(std::string("changcheng")); 63 #endif 64 #if defined(KILL_DEAD) 65 while(1) { 66 #endif 67 int r = image.write(0, bw.length(), bw); 68 #if defined(KILL_DEAD) 69 } 70 #endif 71 72 ceph::bufferlist br; 73 int read = image.read(0, bw.length(), br); 74 br.append(std::string(4, '\0')); 75 std::cout << br.c_str() << std::endl; 76 77 #if defined(KILL_DEAD) 78 while(1); 79 #else 80 sleep(10); 81 #endif 82 83 done: 84 image.close(); 85 io_ctx.close(); 86 rados.shutdown(); 87 exit(EXIT_SUCCESS); 88 } B.R. Changcheng
On Wed, Jul 22, 2020 at 7:46 AM Liu, Changcheng <changcheng.liu@intel.com> wrote:
Hi all, I've checked below document: https://docs.ceph.com/docs/master/rbd/rbd-exclusive-locks/
The content is contradictory to the result shown in below experiment: The experient shows that another process could still write data to rbd volume while there already has a process write to the same rbd volume continuously.
This is the expected behaviour. Exclusive lock is a cooperative mechanism that ensures that only a single client is able to write to the image and update its metadata (such as the object map) at any given moment, not until the client exits. It is acquired automatically and the ownership is transparently transitioned between clients. In your example, "second" wakes up and requests the lock from "first", "first" releases it, "second" performs its write, "first" reacquires the lock and goes on. If you want to disable transparent lock transitions, you need to acquire the lock manually with RBD_LOCK_MODE_EXCLUSIVE:
1. ceph master head commit: commit 1dd932a8f565dc74cf2441ab139aa173575c0e92 Date: Thu Jul 16 09:42:30 2020 +0900
2. setup cluster build$ OSD=3 MON=1 MGR=1 RGW=0 MDS=0 ../src/vstart.sh -k -d -n
3. create rbd volume build$ bin/ceph -c ceph.conf osd pool create rbd 128 build$ bin/ceph -c ceph.conf osd pool application enable rbd rbd build$ bin/rbd -c ceph.conf create fio_test --size 10G --image-format=2 --rbd_default_features=13
4. set environment variable build$ export LD_LIBRARY_PATH=/home/nstcc3/work/src/ceph/build/lib:$LD_LIBRARY_PATH
5. build attached file build$ g++ librbdtest.cpp -DKILL_DEAD -I../src/include -L lib/ -lrados -lrbd -o first build$ g++ librbdtest.cpp -I../src/include -L lib/ -lrados -lrbd -o second
6. run program 1) build$ ./first "first" process write to rbd volume continuously. 2) build$ ./second "second" process could still write to the same volume and exit normally.
source code file: librbdtest.cpp 1 #include <rbd/librbd.hpp> 2 #include <rados/librados.hpp> 3 4 #include <cstring> 5 #include <iostream> 6 #include <string> 7 #include <unistd.h> 8 9 void err_msg(int ret, const std::string &msg = "") { 10 std::cerr << "[error] msg:" << msg << " strerror: " 11 << strerror(-ret) << std::endl; 12 } 13 void err_exit(int ret, const std::string &msg = "") { 14 err_msg(ret, msg); 15 exit(EXIT_FAILURE); 16 } 17 18 int main(int argc, char* argv[]) { 19 #if !defined(KILL_DEAD) 20 sleep(30); 21 #endif 22 int ret = 0; 23 librados::Rados rados; 24 25 ret = rados.init("admin"); 26 if (ret < 0) 27 err_exit(ret,"failed to initialize rados"); 28 ret = rados.conf_read_file("ceph.conf"); 29 if (ret < 0) 30 err_exit(ret, "failed to parse ceph.conf"); 31 32 ret = rados.connect(); 33 if (ret < 0) 34 err_exit(ret, "failed to connect to rados cluster"); 35 36 librados::IoCtx io_ctx; 37 std::string pool_name = "rbd"; 38 ret = rados.ioctx_create(pool_name.c_str(), io_ctx); 39 if (ret < 0) { 40 rados.shutdown(); 41 err_exit(ret, "failed to create ioctx"); 42 } 43 44 // rbd 45 librbd::RBD rbd; 46 47 librbd::Image image; 48 std::string image_name = "fio_test"; 49 ret = rbd.open(io_ctx, image, image_name.c_str()); 50 if (ret < 0) { 51 io_ctx.close(); 52 rados.shutdown(); 53 err_exit(ret, "failed to open rbd image"); 54 } else { 55 std::cout << "open image succeeded" << std::endl; 56 } 57
image.lock_acquire(RBD_LOCK_MODE_EXCLUSIVE); Thanks, Ilya
On 11:10 Wed 22 Jul, Ilya Dryomov wrote:
On Wed, Jul 22, 2020 at 7:46 AM Liu, Changcheng <changcheng.liu@intel.com> wrote:
Hi all, I've checked below document: https://docs.ceph.com/docs/master/rbd/rbd-exclusive-locks/
The content is contradictory to the result shown in below experiment: The experient shows that another process could still write data to rbd volume while there already has a process write to the same rbd volume continuously.
This is the expected behaviour. Exclusive lock is a cooperative mechanism that ensures that only a single client is able to write to the image and update its metadata (such as the object map) at any given moment, not until the client exits. It is acquired automatically and the ownership is transparently transitioned between clients. In your example, "second" wakes up and requests the lock from "first", "first" releases it, "second" performs its write, "first" reacquires the lock and goes on.
If you want to disable transparent lock transitions, you need to acquire the lock manually with RBD_LOCK_MODE_EXCLUSIVE: @Ilya: Thanks for your info. The transparent lock transition could be disabled by acquring the lock with RBD_LOCK_MODE_EXCLUSIVE.
After "first" process acquire the lock with "RBD_LOCK_MODE_EXCLUSIVE", is it possible for another process to be notified that the lock is released whatever the "first" process exit gracefully or be killed? I write below program to run "another process". If I manully remove the lock, "another process" could be notified. However, if the "first" process exit gracefully or be killed, "another process"'s handle_notify won't be called at all. 1 #include <rbd/librbd.hpp> 2 #include <rados/librados.hpp> 3 4 #include <cstring> 5 #include <iostream> 6 #include <string> 7 8 class TestWatcher { 9 public: 10 librados::Rados rados; 11 librbd::RBD rbd; 12 librbd::Image image; 13 librados::IoCtx io_ctx; 14 15 std::string pool_name; 16 std::string image_name; 17 18 TestWatcher(std::string pool_name = "rbd", 19 std::string image_name = "fio_test") 20 : pool_name(pool_name), image_name(image_name) { 21 int ret = rados.init("admin"); 22 if (ret < 0) { 23 std::cout << "failed to initialize rados" << std::endl; 24 exit(1); 25 } 26 27 ret = rados.conf_read_file("ceph.conf"); 28 if (ret < 0) { 29 std::cout << "failed to parse ceph.conf" << std::endl; 30 exit(1); 31 } 32 33 ret = rados.connect(); 34 if (ret < 0) { 35 std::cout << "failed to connect to rados cluster" << std::endl; 36 exit(1); 37 } 38 39 ret = rados.ioctx_create(pool_name.c_str(), io_ctx); 40 if (ret < 0) { 41 rados.shutdown(); 42 std::cout << "failed to create ioctx" << std::endl; 43 exit(1); 44 } 45 46 ret = rbd.open(io_ctx, image, image_name.c_str()); 47 if (ret < 0) { 48 io_ctx.close(); 49 rados.shutdown(); 50 std::cout << "failed to open rbd image" << std::endl; 51 exit(1); 52 } else { 53 std::cout << "open image succeeded" << std::endl; 54 } 55 } 56 57 ~TestWatcher() { 58 image.close(); 59 io_ctx.close(); 60 rados.shutdown(); 61 62 if (watch_ctx != nullptr) { 63 delete watch_ctx; 64 watch_ctx = nullptr; 65 } 66 } 67 68 class WatchCtx: public librbd::UpdateWatchCtx { 69 private: 70 TestWatcher &_test_watcher; 71 public: 72 explicit WatchCtx(TestWatcher &test_watcher): 73 _test_watcher(test_watcher) { 74 } 75 76 int list_watchers() { 77 std::list<librbd::image_watcher_t> watcher_list; 78 int r = _test_watcher.image.list_watchers(watcher_list); 79 if (r >= 0) { 80 for (auto it = watcher_list.cbegin(); it != watcher_list.cend(); ++it) { 81 std::cout << "addr: " << it->addr.c_str() << ", " 82 << "id: " << it->id << ", " 83 << "cookie: " << it->cookie << std::endl; 84 } 85 } 86 return r; 87 } 88 89 int list_lockers() { 90 std::list<librbd::locker_t> lockers; 91 std::string tag; 92 bool exclusive; 93 int r = _test_watcher.image.list_lockers(&lockers, &exclusive, &tag); 94 if (r >= 0) { 95 for (auto it = lockers.cbegin(); it != lockers.cend(); ++it) { 96 std::cout << "client: " << it->client.c_str() << ", " 97 << "cookie: " << it->cookie.c_str() << ", " 98 << "address: " << it->address.c_str() << std::endl; 99 } 100 } 101 return r; 102 } 103 void handle_notify() override { 104 std::cout << "event comming" << std::endl; 105 } 106 }; 107 108 int list_watchers() { 109 watch_ctx = new WatchCtx(*this); 110 int r = watch_ctx->list_watchers(); 111 delete watch_ctx; 112 watch_ctx = nullptr; 113 return r; 114 } 115 116 int list_lockers() { 117 watch_ctx = new WatchCtx(*this); 118 int r = watch_ctx->list_lockers(); 119 delete watch_ctx; 120 watch_ctx = nullptr; 121 return r; 122 } 123 124 125 int update_watch(uint64_t *phandle) { 126 watch_ctx = new WatchCtx(*this); 127 return image.update_watch(watch_ctx, phandle); 128 } 129 WatchCtx* watch_ctx = nullptr; 130 }; 131 132 int main(void) { 133 TestWatcher testwatcher; 134 testwatcher.list_watchers(); 135 testwatcher.list_lockers(); 136 uint64_t handle = 0; 137 testwatcher.update_watch(&handle); 138 while(1); 139 } B.R. Changcheng
Thanks,
Ilya
On Wed, Jul 22, 2020 at 12:34 PM Liu, Changcheng <changcheng.liu@intel.com> wrote:
On 11:10 Wed 22 Jul, Ilya Dryomov wrote:
On Wed, Jul 22, 2020 at 7:46 AM Liu, Changcheng <changcheng.liu@intel.com> wrote:
Hi all, I've checked below document: https://docs.ceph.com/docs/master/rbd/rbd-exclusive-locks/
The content is contradictory to the result shown in below experiment: The experient shows that another process could still write data to rbd volume while there already has a process write to the same rbd volume continuously.
This is the expected behaviour. Exclusive lock is a cooperative mechanism that ensures that only a single client is able to write to the image and update its metadata (such as the object map) at any given moment, not until the client exits. It is acquired automatically and the ownership is transparently transitioned between clients. In your example, "second" wakes up and requests the lock from "first", "first" releases it, "second" performs its write, "first" reacquires the lock and goes on.
If you want to disable transparent lock transitions, you need to acquire the lock manually with RBD_LOCK_MODE_EXCLUSIVE: @Ilya: Thanks for your info. The transparent lock transition could be disabled by acquring the lock with RBD_LOCK_MODE_EXCLUSIVE.
After "first" process acquire the lock with "RBD_LOCK_MODE_EXCLUSIVE", is it possible for another process to be notified that the lock is released whatever the "first" process exit gracefully or be killed?
No. Theoretically, "second" could block, waiting for the lock to be released by "first" (whether gracefully or not), but I don't think librbd does that. (And if it did, it would have been based on periodic retries, not notifies, because if the process is killed there is nowhere for that notify to come from.)
I write below program to run "another process". If I manully remove the lock, "another process" could be notified. However, if the "first" process exit gracefully or be killed, "another process"'s handle_notify won't be called at all.
If you are going to use exclusive lock API, you shouldn't be poking at the underlying watches and notifies.
1 #include <rbd/librbd.hpp> 2 #include <rados/librados.hpp> 3 4 #include <cstring> 5 #include <iostream> 6 #include <string> 7 8 class TestWatcher { 9 public: 10 librados::Rados rados; 11 librbd::RBD rbd; 12 librbd::Image image; 13 librados::IoCtx io_ctx; 14 15 std::string pool_name; 16 std::string image_name; 17 18 TestWatcher(std::string pool_name = "rbd", 19 std::string image_name = "fio_test") 20 : pool_name(pool_name), image_name(image_name) { 21 int ret = rados.init("admin"); 22 if (ret < 0) { 23 std::cout << "failed to initialize rados" << std::endl; 24 exit(1); 25 } 26 27 ret = rados.conf_read_file("ceph.conf"); 28 if (ret < 0) { 29 std::cout << "failed to parse ceph.conf" << std::endl; 30 exit(1); 31 } 32 33 ret = rados.connect(); 34 if (ret < 0) { 35 std::cout << "failed to connect to rados cluster" << std::endl; 36 exit(1); 37 } 38 39 ret = rados.ioctx_create(pool_name.c_str(), io_ctx); 40 if (ret < 0) { 41 rados.shutdown(); 42 std::cout << "failed to create ioctx" << std::endl; 43 exit(1); 44 } 45 46 ret = rbd.open(io_ctx, image, image_name.c_str()); 47 if (ret < 0) { 48 io_ctx.close(); 49 rados.shutdown(); 50 std::cout << "failed to open rbd image" << std::endl; 51 exit(1); 52 } else { 53 std::cout << "open image succeeded" << std::endl; 54 } 55 } 56 57 ~TestWatcher() { 58 image.close(); 59 io_ctx.close(); 60 rados.shutdown(); 61 62 if (watch_ctx != nullptr) { 63 delete watch_ctx; 64 watch_ctx = nullptr; 65 } 66 } 67 68 class WatchCtx: public librbd::UpdateWatchCtx { 69 private: 70 TestWatcher &_test_watcher; 71 public: 72 explicit WatchCtx(TestWatcher &test_watcher): 73 _test_watcher(test_watcher) { 74 } 75 76 int list_watchers() { 77 std::list<librbd::image_watcher_t> watcher_list; 78 int r = _test_watcher.image.list_watchers(watcher_list); 79 if (r >= 0) { 80 for (auto it = watcher_list.cbegin(); it != watcher_list.cend(); ++it) { 81 std::cout << "addr: " << it->addr.c_str() << ", " 82 << "id: " << it->id << ", " 83 << "cookie: " << it->cookie << std::endl; 84 } 85 } 86 return r; 87 } 88 89 int list_lockers() { 90 std::list<librbd::locker_t> lockers; 91 std::string tag; 92 bool exclusive; 93 int r = _test_watcher.image.list_lockers(&lockers, &exclusive, &tag); 94 if (r >= 0) { 95 for (auto it = lockers.cbegin(); it != lockers.cend(); ++it) { 96 std::cout << "client: " << it->client.c_str() << ", " 97 << "cookie: " << it->cookie.c_str() << ", " 98 << "address: " << it->address.c_str() << std::endl; 99 } 100 } 101 return r; 102 } 103 void handle_notify() override { 104 std::cout << "event comming" << std::endl; 105 } 106 }; 107 108 int list_watchers() { 109 watch_ctx = new WatchCtx(*this); 110 int r = watch_ctx->list_watchers(); 111 delete watch_ctx; 112 watch_ctx = nullptr; 113 return r; 114 } 115 116 int list_lockers() { 117 watch_ctx = new WatchCtx(*this); 118 int r = watch_ctx->list_lockers(); 119 delete watch_ctx; 120 watch_ctx = nullptr; 121 return r; 122 } 123 124 125 int update_watch(uint64_t *phandle) { 126 watch_ctx = new WatchCtx(*this); 127 return image.update_watch(watch_ctx, phandle); 128 } 129 WatchCtx* watch_ctx = nullptr; 130 }; 131 132 int main(void) { 133 TestWatcher testwatcher; 134 testwatcher.list_watchers(); 135 testwatcher.list_lockers(); 136 uint64_t handle = 0; 137 testwatcher.update_watch(&handle); 138 while(1); 139 }
Thanks, Ilya
On 13:44 Wed 22 Jul, Ilya Dryomov wrote:
On Wed, Jul 22, 2020 at 12:34 PM Liu, Changcheng <changcheng.liu@intel.com> wrote:
On 11:10 Wed 22 Jul, Ilya Dryomov wrote:
On Wed, Jul 22, 2020 at 7:46 AM Liu, Changcheng <changcheng.liu@intel.com> wrote:
If you want to disable transparent lock transitions, you need to acquire the lock manually with RBD_LOCK_MODE_EXCLUSIVE: @Ilya: Thanks for your info. The transparent lock transition could be disabled by acquring the lock with RBD_LOCK_MODE_EXCLUSIVE.
After "first" process acquire the lock with "RBD_LOCK_MODE_EXCLUSIVE", is it possible for another process to be notified that the lock is released whatever the "first" process exit gracefully or be killed?
No. Theoretically, "second" could block, waiting for the lock to be released by "first" (whether gracefully or not), but I don't think librbd does that. (And if it did, it would have been based on periodic retries, not notifies, because if the process is killed there is nowhere for that notify to come from.)
@Ilya: The curious thing is that the "second" process could be notified if I remove the lock manually i.e. use "rbd lock rm" command. --Thanks Changcheng
participants (2)
-
Ilya Dryomov
-
Liu, Changcheng