Comparison of 3 replication models on Pech OSD cluster
Hi all, I would like to share the comparison of 3 replication models based on Pech OSD [1] cluster, which supports a sufficient minimum to replicate transactions from OSD to OSD and keeps all data mutations in memory (memstore). My goal was to compare "primary-copy", "chain" and "client-based" replication models and answer the question: how each model affects network performance. For this estimation I chose to implement my own OSD with bare minimum (laborious but worth it) which design is similar to Crimson OSD but core is based on sources from kernel libceph implementation (i.e. messenger, osdmap, mon_client, etc), thus written in pure C. -- What Pech OSD supports and what does not -- Comparison of the network response under different replication scenarios does not require fail-over (we assume that during testing, storing data in memory never fails, hosts never crash, etc), thus to ease development Pech OSD does not support (current state of the code) peering and fail-over. So object modification is replicated on each mutation, yes, but cluster is not able to come to the consistent state after an error. Pech OSD supports RBD images, so that image can be accessed from userspace librbd or mapped by kernel RBD block device. That is a bare minimum which I need to run FIO loads and test network behavior. -- What I test -- Originally my goal was to compare performance under same loads but using different replication models: "client-based", "primary-copy" and "chain". I want to see the numbers what different models can bring in terms of network bandwidth, latency and IOPS (and when we talk about comparison of replication models, only network is the factor which impacts the overall performance). Shortly about replication models: "client-based" - client itself is responsible for sending requests to replicas. To test this model OSD client code was modified on userspace [2] and kernel [3] sides. Pros: savings on network hops which reduces latency. Cons: complications in replication algorithm when PG is not healthy, complications in replication algorithm when there is a concurrent access to the same object from divers clients, client network should be fat enough. "chain" - client sends write request to primary, primary forwards to the next secondary, and so on. Final ACK from last replica in chain reaches primary or client directly. Pros: each OSD sends a request only once, which reduces load on network for particular node and spreads load. Overall bandwidth should increase. Cons: sequential requests processing, which should impacts latency. "primary-copy" - default and the only one model for Ceph: client accesses primary replica, primary replica fans out data to secondaries. Pros: already implemented. Cons: higher latency comparing to "client-based", lower bandwidth comparing to "chain". What is said above is the theory which has motivated me to prove or disprove it with numbers on the real cluster. -- How I test -- I have the cluster at my disposal, with 5 hosts with 100gb network for OSDs and 8 client hosts, with 25gbit/s network. Each OSD host has 24 CPUs, so for obvious reasons each host runs 24 OSDs, so (24x5) 120 Pech OSDs for the whole cluster setup. There is one fully declustered pool with 1024 PGs (I want to spread the load as much as possible). Pool is created with 3x replication factor. Each client starts 16 FIO jobs with random write to 16 RBD images (userspace RBD client) with various block sizes, i.e. one FIO job per image and 128 (16x8) jobs in total. Each client host runs FIO server, all data from all servers is aggregated by FIO client and stored in json format. There is a convenient python script [4] which generates and runs FIO jobs, parses json results and outputs them in a human readable pretty table. Major FIO options: ioengine=rbd clientname=admin pool=rbd rw=randwrite size=256m time_based=1 runtime=10 ramp_time=10 iodepth=32 numjobs=1 During all tests I collected almost 1Gb of json data results. Pretty enough for good analysis. -- Results -- Firstly I would like to start comparing "primary-copy" and "chain" on Pech OSD: 120OSDS/pech/primary-copy write/iops write/bw write/clat_ns/mean 4k 365.89 K 1.40 GB/s 11.11 ms 8k 330.51 K 2.52 GB/s 12.22 ms 16k 274.06 K 4.19 GB/s 14.79 ms 32k 204.36 K 6.25 GB/s 19.95 ms 64k 141.78 K 8.68 GB/s 28.54 ms 128k 70.42 K 8.64 GB/s 58.99 ms 256k 37.75 K 9.30 GB/s 109.75 ms 512k 17.46 K 8.67 GB/s 216.53 ms 1m 8.56 K 8.65 GB/s 474.94 ms 120OSDS/pech/chain write/iops write/bw write/clat_ns/mean 4k 380.29 K 1.45 GB/s 10.72 ms 8k 339.10 K 2.59 GB/s 11.99 ms 16k 280.28 K 4.28 GB/s 14.34 ms 32k 206.84 K 6.32 GB/s 19.64 ms 64k 131.57 K 8.05 GB/s 30.54 ms 128k 74.78 K 9.18 GB/s 54.25 ms 256k 39.82 K 9.81 GB/s 103.27 ms 512k 18.47 K 9.17 GB/s 213.78 ms 1m 8.98 K 9.08 GB/s 461.12 ms There is a slight difference in the direction of bandwidth increase for "chain" model, but I would rather take it for a noise. Another runs for similar configuration show almost similar results: there is a minor "bandwidth" improve but not so solid. Client-based results are much more interesting: 120OSDS/pech/client-based write/iops write/bw write/clat_ns/mean 4k 534.08 K 2.04 GB/s 7.62 ms 8k 471.78 K 3.60 GB/s 8.64 ms 16k 367.12 K 5.61 GB/s 11.11 ms 32k 242.56 K 7.41 GB/s 16.82 ms 64k 124.54 K 7.63 GB/s 32.98 ms 128k 62.45 K 7.67 GB/s 66.71 ms 256k 31.10 K 7.69 GB/s 135.36 ms 512k 15.41 K 7.71 GB/s 282.41 ms 1m 7.63 K 7.82 GB/s 567.63 ms Small blocks show significant improve in latency: almost 40%, from 380k IOPS to 534k IOPS. Starting from 64k block the client network 25gbit/s is reached ("client-based" replication means client is responsible for sending the data to all replicas, that means that each byte with 3x replication factor should be repeated 3 times from each client host, having ~8GB/s for 8 clients we estimate each client sends ~1GB/s, with 3x replication factor this is ~3GB/s and this is exactly the ~24gbit/s of the client network). What is important to keep in mind with Pech OSD design is that each OSD process has only 1 OS thread, so when request is received and request handler is executed there is no any preemption happens and no other requests can be handled in parallel (unless special scheduling routine is called, which is not, at least in current code state). So various PGs on particular Pech OSD are handled sequentially. The design is highly CPU bound, thus one simple trick can be made to increase bandwidth: OSD pinning to CPU. Since we have 24 OSDs and 24 CPUs CPU affinitty is easy to apply: 120OSDS-AFF/pech/primary-copy write/iops write/bw write/clat_ns/mean 4k 324.15 K 1.24 GB/s 12.35 ms 8k 293.52 K 2.24 GB/s 13.43 ms 16k 235.53 K 3.60 GB/s 16.46 ms 32k 187.31 K 5.73 GB/s 20.77 ms 64k 170.60 K 10.43 GB/s 23.10 ms 128k 92.54 K 11.33 GB/s 34.48 ms 256k 47.69 K 11.73 GB/s 97.32 ms 512k 18.52 K 9.19 GB/s 252.26 ms 1m 9.20 K 9.28 GB/s 507.33 ms Bandwidth looks better for bigger blocks. In conclusion about replication models. I did not notice any significant difference between "primary-copy" and "chain". Perhaps it makes sense to play with the replication factor. In its turn "client-based" replication can be very promising for loads in homogeneous networks, where there is no any concurrent access to images. Simple example is a cluster with compute and storage nodes in private network, where VMs access their own images. For such setups latency is a factor which plays a huge role. -- Roman [1] https://github.com/rouming/pech [2] https://github.com/rouming/ceph/tree/pech-osd [3] https://github.com/rouming/linux/tree/akpm--ceph-client-based-replication [4] https://github.com/rouming/pech/blob/master/scripts/fio-runner.py
Hi Roman, It's always really interesting to read your messages :) maybe you'll join our telegram chat @ceph_ru? One of your colleagues is there :) Client-based replication is of course the fastest, but the problem is that it's unclear how to provide consistency with it. Maybe it's possible with some restrictions, but... I don't think it's possible in Ceph :) By the way, have you tested their Crimson OSD? Is it any faster than current implementation? (regarding iodepth=1 fsync=1 latency)
Hi all,
I would like to share the comparison of 3 replication models based on Pech OSD [1] cluster, which supports a sufficient minimum to replicate transactions from OSD to OSD and keeps all data mutations in memory (memstore).
My goal was to compare "primary-copy", "chain" and "client-based" replication models and answer the question: how each model affects network performance.
For this estimation I chose to implement my own OSD with bare minimum (laborious but worth it) which design is similar to Crimson OSD but core is based on sources from kernel libceph implementation (i.e. messenger, osdmap, mon_client, etc), thus written in pure C.
-- What Pech OSD supports and what does not --
Comparison of the network response under different replication scenarios does not require fail-over (we assume that during testing, storing data in memory never fails, hosts never crash, etc), thus to ease development Pech OSD does not support (current state of the code) peering and fail-over. So object modification is replicated on each mutation, yes, but cluster is not able to come to the consistent state after an error.
Pech OSD supports RBD images, so that image can be accessed from userspace librbd or mapped by kernel RBD block device. That is a bare minimum which I need to run FIO loads and test network behavior.
-- What I test --
Originally my goal was to compare performance under same loads but using different replication models: "client-based", "primary-copy" and "chain". I want to see the numbers what different models can bring in terms of network bandwidth, latency and IOPS (and when we talk about comparison of replication models, only network is the factor which impacts the overall performance).
Shortly about replication models:
"client-based" - client itself is responsible for sending requests to replicas. To test this model OSD client code was modified on userspace [2] and kernel [3] sides. Pros: savings on network hops which reduces latency. Cons: complications in replication algorithm when PG is not healthy, complications in replication algorithm when there is a concurrent access to the same object from divers clients, client network should be fat enough.
"chain" - client sends write request to primary, primary forwards to the next secondary, and so on. Final ACK from last replica in chain reaches primary or client directly. Pros: each OSD sends a request only once, which reduces load on network for particular node and spreads load. Overall bandwidth should increase. Cons: sequential requests processing, which should impacts latency.
"primary-copy" - default and the only one model for Ceph: client accesses primary replica, primary replica fans out data to secondaries. Pros: already implemented. Cons: higher latency comparing to "client-based", lower bandwidth comparing to "chain".
What is said above is the theory which has motivated me to prove or disprove it with numbers on the real cluster.
-- How I test --
I have the cluster at my disposal, with 5 hosts with 100gb network for OSDs and 8 client hosts, with 25gbit/s network.
Each OSD host has 24 CPUs, so for obvious reasons each host runs 24 OSDs, so (24x5) 120 Pech OSDs for the whole cluster setup.
There is one fully declustered pool with 1024 PGs (I want to spread the load as much as possible). Pool is created with 3x replication factor.
Each client starts 16 FIO jobs with random write to 16 RBD images (userspace RBD client) with various block sizes, i.e. one FIO job per image and 128 (16x8) jobs in total. Each client host runs FIO server, all data from all servers is aggregated by FIO client and stored in json format. There is a convenient python script [4] which generates and runs FIO jobs, parses json results and outputs them in a human readable pretty table.
Major FIO options:
ioengine=rbd clientname=admin pool=rbd
rw=randwrite size=256m
time_based=1 runtime=10 ramp_time=10
iodepth=32 numjobs=1
During all tests I collected almost 1Gb of json data results. Pretty enough for good analysis.
-- Results --
Firstly I would like to start comparing "primary-copy" and "chain" on Pech OSD:
120OSDS/pech/primary-copy
write/iops write/bw write/clat_ns/mean 4k 365.89 K 1.40 GB/s 11.11 ms 8k 330.51 K 2.52 GB/s 12.22 ms 16k 274.06 K 4.19 GB/s 14.79 ms 32k 204.36 K 6.25 GB/s 19.95 ms 64k 141.78 K 8.68 GB/s 28.54 ms 128k 70.42 K 8.64 GB/s 58.99 ms 256k 37.75 K 9.30 GB/s 109.75 ms 512k 17.46 K 8.67 GB/s 216.53 ms 1m 8.56 K 8.65 GB/s 474.94 ms
120OSDS/pech/chain
write/iops write/bw write/clat_ns/mean 4k 380.29 K 1.45 GB/s 10.72 ms 8k 339.10 K 2.59 GB/s 11.99 ms 16k 280.28 K 4.28 GB/s 14.34 ms 32k 206.84 K 6.32 GB/s 19.64 ms 64k 131.57 K 8.05 GB/s 30.54 ms 128k 74.78 K 9.18 GB/s 54.25 ms 256k 39.82 K 9.81 GB/s 103.27 ms 512k 18.47 K 9.17 GB/s 213.78 ms 1m 8.98 K 9.08 GB/s 461.12 ms
There is a slight difference in the direction of bandwidth increase for "chain" model, but I would rather take it for a noise. Another runs for similar configuration show almost similar results: there is a minor "bandwidth" improve but not so solid.
Client-based results are much more interesting:
120OSDS/pech/client-based
write/iops write/bw write/clat_ns/mean 4k 534.08 K 2.04 GB/s 7.62 ms 8k 471.78 K 3.60 GB/s 8.64 ms 16k 367.12 K 5.61 GB/s 11.11 ms 32k 242.56 K 7.41 GB/s 16.82 ms 64k 124.54 K 7.63 GB/s 32.98 ms 128k 62.45 K 7.67 GB/s 66.71 ms 256k 31.10 K 7.69 GB/s 135.36 ms 512k 15.41 K 7.71 GB/s 282.41 ms 1m 7.63 K 7.82 GB/s 567.63 ms
Small blocks show significant improve in latency: almost 40%, from 380k IOPS to 534k IOPS. Starting from 64k block the client network 25gbit/s is reached ("client-based" replication means client is responsible for sending the data to all replicas, that means that each byte with 3x replication factor should be repeated 3 times from each client host, having ~8GB/s for 8 clients we estimate each client sends ~1GB/s, with 3x replication factor this is ~3GB/s and this is exactly the ~24gbit/s of the client network).
What is important to keep in mind with Pech OSD design is that each OSD process has only 1 OS thread, so when request is received and request handler is executed there is no any preemption happens and no other requests can be handled in parallel (unless special scheduling routine is called, which is not, at least in current code state). So various PGs on particular Pech OSD are handled sequentially.
The design is highly CPU bound, thus one simple trick can be made to increase bandwidth: OSD pinning to CPU. Since we have 24 OSDs and 24 CPUs CPU affinitty is easy to apply:
120OSDS-AFF/pech/primary-copy
write/iops write/bw write/clat_ns/mean 4k 324.15 K 1.24 GB/s 12.35 ms 8k 293.52 K 2.24 GB/s 13.43 ms 16k 235.53 K 3.60 GB/s 16.46 ms 32k 187.31 K 5.73 GB/s 20.77 ms 64k 170.60 K 10.43 GB/s 23.10 ms 128k 92.54 K 11.33 GB/s 34.48 ms 256k 47.69 K 11.73 GB/s 97.32 ms 512k 18.52 K 9.19 GB/s 252.26 ms 1m 9.20 K 9.28 GB/s 507.33 ms
Bandwidth looks better for bigger blocks.
In conclusion about replication models. I did not notice any significant difference between "primary-copy" and "chain". Perhaps it makes sense to play with the replication factor.
In its turn "client-based" replication can be very promising for loads in homogeneous networks, where there is no any concurrent access to images. Simple example is a cluster with compute and storage nodes in private network, where VMs access their own images. For such setups latency is a factor which plays a huge role.
-- Roman
[1] https://github.com/rouming/pech [2] https://github.com/rouming/ceph/tree/pech-osd [3] https://github.com/rouming/linux/tree/akpm--ceph-client-based-replication [4] https://github.com/rouming/pech/blob/master/scripts/fio-runner.py _______________________________________________ Dev mailing list -- dev@ceph.io To unsubscribe send an email to dev-leave@ceph.io
Hi Vitali, Sorry for a long response, I was on vacation. On 2020-07-20 00:44, vitalif@yourcmc.ru wrote:
Hi Roman,
It's always really interesting to read your messages :) maybe you'll join our telegram chat @ceph_ru? One of your colleagues is there :)
Promise a lot of fun? ;)
Client-based replication is of course the fastest,
With this comparison I answer the question: if it is fastest, then how much. Because obvious "of course is the fastest", well, not so reasoned :)
but the problem is that it's unclear how to provide consistency with it.
Here I tried to highlight client-based replication problems: https://lists.ceph.io/hyperkitty/list/dev@ceph.io/thread/N46NR7NBHWBQL4B2ASU... And yes, comes with restrictions, but there are scenarios which do not require strong sequential consistency of log-based replication, e.g. if you run an 1 rbd client per 1 image and a filesystem on top with journaling and strong requests order why not to rely on the filesystem recovery mechanisms? That is the question which bothers me for quite a while and that is exactly the reason why I started pech osd: to find some answers.
Maybe it's possible with some restrictions, but... I don't think it's possible in Ceph :)
No, for sure not. Not RADOS strong consistency semantics. But why? Take what you need and cut what is useless to fit your requirements.
By the way, have you tested their Crimson OSD?
Yes, I did. But with Crimson OSD everything went wrong: the first problem I came across is that I was not able to reach desired number of 120 OSDs running: the average number after various restarts I got from monitor is ~50. I did not try to debug and simply reduced number of OSDs to 35 (5 hosts, 7 ODSs on each) and reran all tests, so here is the results for all types of osds (for fair reference): 35OSDS/crimson/primary-copy write/iops write/bw write/clat_ns/mean 4k 3.50 K 14.30 MB/s 888.00 ms 8k 4.97 K 40.37 MB/s 687.27 ms 16k 4.90 K 79.63 MB/s 709.63 ms 32k 4.48 K 145.50 MB/s 703.80 ms 64k 4.46 K 290.72 MB/s 731.06 ms 128k 4.38 K 570.44 MB/s 720.32 ms 256k 4.13 K 1.05 GB/s 755.77 ms 512k 2.56 K 1.32 GB/s 1.15 s 1m 1.16 K 1.24 GB/s 2.22 s 35OSDS/ceph/primary-copy write/iops write/bw write/clat_ns/mean 4k 90.74 K 355.92 MB/s 44.51 ms 8k 75.03 K 588.98 MB/s 53.45 ms 16k 92.58 K 1.42 GB/s 42.10 ms 32k 122.95 K 3.76 GB/s 33.27 ms 64k 101.28 K 6.20 GB/s 39.45 ms 128k 45.28 K 5.57 GB/s 83.26 ms 256k 26.08 K 6.44 GB/s 138.16 ms 512k 14.49 K 7.23 GB/s 254.74 ms 1m 5.91 K 6.05 GB/s 588.28 ms 35OSDS/pech/primary-copy write/iops write/bw write/clat_ns/mean 4k 289.22 K 1.10 GB/s 14.94 ms 8k 231.93 K 1.77 GB/s 15.94 ms 16k 228.60 K 3.49 GB/s 17.28 ms 32k 208.95 K 6.39 GB/s 19.08 ms 64k 106.66 K 6.53 GB/s 37.69 ms 128k 53.48 K 6.57 GB/s 73.03 ms 256k 25.03 K 6.19 GB/s 139.59 ms 512k 12.63 K 6.32 GB/s 302.50 ms 1m 5.91 K 6.05 GB/s 650.03 ms I did not notice anything strange in crimsons logs and did try to debug, so do not know why the results are so bad for the crimson case.
Is it any faster than current implementation? (regarding iodepth=1 fsync=1 latency)
My original goal was to test real distributed load: many osd hosts, many clients hosts (I was keen to see how Pech was behaving). Your "latency" load does not require a cluster setup and can be executed on a localhost with 3 osds (x3 replication), so here are the results: "-o ms_crc_data=false -o debug_osd=0 -o debug_ms=0" rbd.fio rw=randwrite iodepth=1 numjobs=1 runtume=10 size=256m /// crimson-osd 4k IOPS=101, BW=406KiB/s, Lat=9846.09usec 8k IOPS=100, BW=802KiB/s, Lat=9973.39usec 16k IOPS=99, BW=1599KiB/s, Lat=10000.17usec 32k IOPS=96, BW=3088KiB/s, Lat=10355.56usec 64k IOPS=591, BW=36.0MiB/s, Lat=1687.63usec 128k IOPS=508, BW=63.6MiB/s, Lat=1963.95usec 256k IOPS=379, BW=94.9MiB/s, Lat=2632.29usec 512k IOPS=338, BW=169MiB/s, Lat=2952.39usec 1m IOPS=166, BW=166MiB/s, Lat=6011.29usec /// ceph-osd 4k IOPS=1908, BW=7634KiB/s, Lat=522.07usec 8k IOPS=1838, BW=14.4MiB/s, Lat=542.10usec 16k IOPS=1751, BW=27.4MiB/s, Lat=568.98usec 32k IOPS=2048, BW=64.0MiB/s, Lat=486.48usec 64k IOPS=1985, BW=124MiB/s, Lat=501.80usec 128k IOPS=1869, BW=234MiB/s, Lat=532.96usec 256k IOPS=1645, BW=411MiB/s, Lat=605.66usec 512k IOPS=1195, BW=598MiB/s, Lat=833.64usec 1m IOPS=704, BW=705MiB/s, Lat=1414.01usec /// pech-osd OSD=X; CEPH=~/devel/ceph-upstream; ./pech-osd --mon_addrs 192.168.0.97:50001 --server_ip 0.0.0.0 --name $OSD --fsid `cat $CEPH/build/dev/osd$OSD/fsid` --class_dir $CEPH/build/lib --log_level 5 --replication primary-copy --nocrc 4k IOPS=5618, BW=21.9MiB/s, Lat=176.48usec 8k IOPS=5654, BW=44.2MiB/s, Lat=175.26usec 16k IOPS=5504, BW=86.0MiB/s, Lat=180.20usec 32k IOPS=4976, BW=156MiB/s, Lat=199.37usec 64k IOPS=4334, BW=271MiB/s, Lat=229.09usec 128k IOPS=3397, BW=425MiB/s, Lat=292.52usec 256k IOPS=2392, BW=598MiB/s, Lat=416.12usec 512k IOPS=1505, BW=753MiB/s, Lat=661.25usec 1m IOPS=687, BW=688MiB/s, Lat=1446.60usec Results should be treated carefully, since practically these numbers are almost certainly unreachable, but here you are right: numbers give a clear upper bound. -- Roman
Wow, crimson osd has some big problems if its latency is 10ms. OK, let's just wait until it's fixed. Problem is that you don't know which data to resync without journaling. AFAIK Linstor/drbd9 does a similar thing, something like "write intent journaling", to provide fast resyncs, even though they don't have EC (they have it in beta). So I think the client-based approach can only be possible without fully ditching atomicity. Maybe for example it should involve some additional lazy synchronization among OSDs themselves using pglogs from client-driven writes. Hm.. sounds like it's even not unreal in ceph..:-) Regarding the latency test, it also depends on network latency, so local test isn't the same as a clustered one. But thanks anyway. Pech OSD seems fast :-) -- With best regards, Vitaliy Filippov
On 2020-07-29 01:33, Виталий Филиппов wrote:
Wow, crimson osd has some big problems if its latency is 10ms. OK, let's just wait until it's fixed.
I really keen to see good crimson results, since I need a good reference numbers to understand where pech sucks, but each test I do I fail to get something significant.
Problem is that you don't know which data to resync without journaling.
Take simple RAID1 as an example. Nothing can be simpler. It does resync according to the bitmap of dirty blocks (when enabled). No journaling is involved. Why object data mutations (write IOs to an object) can't be tracked using old-school bitmap? Let's avoid any parallel access to the same image, here I consider only one particular case: VM opens its own image, i.e. one image per one client, thus no parallel access (with parallel access everything gets complicated, but IMHO not unsolvable). So having some restrictions we can track object data modifications (i.e. plain write IOs) in a bitmap.
AFAIK Linstor/drbd9 does a similar thing, something like "write intent journaling", to provide fast resyncs, even though they don't have EC (they have it in beta). So I think the client-based approach can only be possible without fully ditching atomicity. Maybe for example it should involve some additional lazy synchronization among OSDs themselves using pglogs from client-driven writes. Hm.. sounds like it's even not unreal in ceph..:-)
Unfortunately I can say nothing regarding linstor, but having simple (what is important to me is plain C :) and fast pech osd implementation of basic logic (transport and monitor client) it is possible to test and prove/disprove some interesting ideas.
Regarding the latency test, it also depends on network latency, so local test isn't the same as a clustered one.
True, is not the same, but we compare osds between each other, not the localhost network and cluster network. Just extrapolate the difference between osds to any other network setup and you will be almost certainly correct :) And frankly, now I do not have any real cluster in my disposal. So could test only on my local machine.
But thanks anyway. Pech OSD seems fast :-)
To reach current crimson state peering has to be implemented, though I never had plans to compete with crimson. What especially warms my heart is compilation time: 4 seconds on 8 cpus :) -- Roman
Why object data mutations (write IOs to an object) can't be tracked using old-school bitmap?
That's what I call "write intent journaling" %) it's still a sort of "journaling" because you first modify the bitmap, fsync it, then proceed with the write itself. Ceph/Pech architecture, however, dictates that an RBD spans multiple OSDs. Because if it doesn't it stops being Ceph and becomes Linstor. :) and there is a Linstor already. :) in this case it seems that a journal is more convenient than a plain bitmap... aaand this is precisely what PGlog is (pglog isn't a journal with data, it only contains a list of updated objects). -- With best regards, Vitaliy Filippov
On 2020-07-30 01:02, vitalif@yourcmc.ru wrote:
Why object data mutations (write IOs to an object) can't be tracked using old-school bitmap?
That's what I call "write intent journaling" %) it's still a sort of "journaling" because you first modify the bitmap, fsync it, then proceed with the write itself.
Not quite. Journal has to be updated on *each* mutation (either data or meta-data). With bitmap you can mark block as dirty once per say N seconds, so if there are a lot of writes to that block you save a lot of fsyncs and just write to the block directly. That increases amount of data to be resynced (blocks still marked as dirty) in case of crash, but that is minor. Also, journal and actual data updates go through transactions to guarantee atomicity (well, everything goes through transactions). You can't have a record in the journal and no data updated, the opposite is also true: you can't have data updated and no record in a journal. Bitmap relaxes this restriction. When you have sequential transactions for the whole pg (pg lock and friends) performance degrades.
Ceph/Pech architecture, however, dictates that an RBD spans multiple OSDs. Because if it doesn't it stops being Ceph and becomes Linstor. :) and there is a Linstor already. :)
Linux does not become Windows, even both draw windows on the screen :)
in this case it seems that a journal is more convenient than a plain bitmap... aaand this is precisely what PGlog is (pglog isn't a journal with data, it only contains a list of updated objects).
True, data updates are not covered by any journal, but on each write IO you have to add a record to the pg journal, replicate journal update with data update and atomically (through transaction) apply updates locally. And these updates in the same pg (even to different objects) are fully serialized, indeed very convenient :) -- Roman
participants (3)
-
Roman Penyaev
-
vitalif@yourcmc.ru
-
Виталий Филиппов