I observed that on an otherwise idle cluster, scrubbing cannot fully utilise the speed of my HDDs. `iostat` shows only 8-10 MB/s per disk, instead of the ~100 MB/s most HDDs can easily deliver. Changing scrubbing settings does not help (see below). Environment: * 6 active+clean+scrubbing+deep * Ceph version 16.2.7. * BlueStore * My cluster has many objects small objects ("402.32M objects, 38 TiB" from "ceph status") due to small files (4 - 32 KiB) on CephFS. `iostat -x 5` with default settings: Device r/s rkB/s rrqm/s %rrqm r_await rareq-sz w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util dm-0 198.60 6878.40 0.00 0.00 12.78 34.63 51.80 2612.80 0.00 0.00 14.82 50.44 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 3.30 91.24 dm-1 0.80 3.20 0.00 0.00 11.50 4.00 52.60 2582.40 0.00 0.00 13.69 49.10 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.73 3.78 dm-10 11.20 71.20 0.00 0.00 0.09 6.36 145.80 583.20 0.00 0.00 0.14 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.02 2.62 dm-11 192.60 6737.60 0.00 0.00 10.74 34.98 34.80 1684.80 0.00 0.00 11.47 48.41 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 2.47 91.40 dm-12 245.40 10194.40 0.00 0.00 9.43 41.54 21.20 575.20 0.00 0.00 3.94 27.13 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 2.40 87.92 dm-13 30.80 1772.80 0.00 0.00 11.61 57.56 78.80 4507.20 0.00 0.00 19.78 57.20 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.92 9.54 dm-14 3.20 24.80 0.00 0.00 0.12 7.75 125.20 500.80 0.00 0.00 0.12 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 2.18 dm-15 2.80 19.20 0.00 0.00 0.14 6.86 105.40 421.60 0.00 0.00 0.05 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 1.76 dm-16 0.80 6.40 0.00 0.00 0.00 8.00 111.00 444.00 0.00 0.00 0.10 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 1.82 dm-17 10.80 76.80 0.00 0.00 0.09 7.11 151.40 605.60 0.00 0.00 0.08 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 2.92 dm-18 10.20 67.20 0.00 0.00 0.08 6.59 115.60 462.40 0.00 0.00 0.04 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 2.16 dm-19 10.20 56.80 0.00 0.00 0.10 5.57 109.00 436.00 0.00 0.00 0.07 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 2.34 dm-2 4.80 435.20 0.00 0.00 0.12 90.67 751.80 6292.80 0.00 0.00 0.07 8.37 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.05 16.14 dm-20 0.40 2.40 0.00 0.00 0.00 6.00 265.00 2459.20 0.00 0.00 0.10 9.28 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.03 5.36 dm-21 191.00 6105.60 0.00 0.00 6.34 31.97 67.80 3748.00 0.00 0.00 19.56 55.28 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 2.54 42.22 dm-3 1.00 8.80 0.00 0.00 0.00 8.80 91.00 364.00 0.00 0.00 0.04 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.54 dm-4 167.60 4973.60 0.00 0.00 10.15 29.68 49.20 2511.20 0.00 0.00 11.39 51.04 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 2.26 89.18 dm-5 11.20 73.60 0.00 0.00 0.12 6.57 124.40 497.60 0.00 0.00 0.07 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 2.16 dm-6 27.20 1644.80 0.00 0.00 12.22 60.47 57.20 3316.80 0.00 0.00 15.93 57.99 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.24 6.78 dm-7 217.40 8032.80 0.00 0.00 12.04 36.95 64.40 3654.40 0.00 0.00 23.69 56.75 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 4.14 97.28 dm-8 10.80 70.40 0.00 0.00 0.15 6.52 111.80 447.20 0.00 0.00 0.04 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 2.32 dm-9 1.60 8.00 0.00 0.00 13.25 5.00 46.60 2563.20 0.00 0.00 9.08 55.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.44 3.70 md127 1.60 107.20 0.00 0.00 0.00 67.00 142.00 1856.00 0.00 0.00 0.05 13.07 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 2.80 nvme0n1 42.00 772.00 0.00 0.00 0.10 18.38 1001.00 10503.50 451.20 31.07 0.03 10.49 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.04 25.20 nvme1n1 37.20 248.00 0.00 0.00 0.09 6.67 609.80 6723.50 369.00 37.70 0.04 11.03 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.03 15.80 sda 0.80 3.20 0.00 0.00 11.50 4.00 33.60 2582.40 19.00 36.12 8.43 76.86 0.00 0.00 0.00 0.00 0.00 0.00 3.20 0.25 0.29 3.78 sdb 117.80 10195.20 127.80 52.04 11.53 86.55 18.80 575.20 2.40 11.32 3.93 30.60 0.00 0.00 0.00 0.00 0.00 0.00 2.80 3.93 1.44 87.90 sdc 4.60 1644.80 22.60 83.09 11.74 357.57 23.60 3316.80 33.60 58.74 8.53 140.54 0.00 0.00 0.00 0.00 0.00 0.00 1.80 4.56 0.26 6.74 sdd 109.40 4975.20 58.40 34.80 12.19 45.48 22.00 2511.20 27.20 55.28 5.75 114.15 0.00 0.00 0.00 0.00 0.00 0.00 2.80 6.07 1.48 89.16 sde 115.40 6563.20 77.00 40.02 12.21 56.87 26.20 1684.80 8.60 24.71 7.82 64.31 0.00 0.00 0.00 0.00 0.00 0.00 1.80 6.78 1.63 91.36 sdf 6.40 1772.80 24.40 79.22 14.84 277.00 39.20 4507.20 39.60 50.25 13.93 114.98 0.00 0.00 0.00 0.00 0.00 0.00 3.20 2.25 0.65 9.54 sdg 121.60 8033.60 95.80 44.07 12.39 66.07 30.60 3654.40 33.80 52.48 12.47 119.42 0.00 0.00 0.00 0.00 0.00 0.00 2.60 8.77 1.91 97.24 sdh 122.00 6105.60 69.00 36.13 5.47 50.05 33.00 3748.00 34.80 51.33 16.94 113.58 0.00 0.00 0.00 0.00 0.00 0.00 2.20 5.00 1.24 42.20 sdi 117.00 6856.80 81.60 41.09 12.25 58.61 32.00 2612.80 19.80 38.22 10.80 81.65 0.00 0.00 0.00 0.00 0.00 0.00 2.60 8.46 1.80 91.18 sdj 1.60 8.00 0.00 0.00 13.12 5.00 31.20 2563.20 15.40 33.05 5.06 82.15 0.00 0.00 0.00 0.00 0.00 0.00 2.00 1.70 0.18 3.70 With settings osd_deep_scrub_stride = 4194304 osd_scrub_load_threshold = 20 osd_scrub_chunk_min = 15 osd_scrub_chunk_max = 75 osd_max_scrubs = 3 I get a slight improvement only: Device r/s rkB/s rrqm/s %rrqm r_await rareq-sz w/s wkB/s wrqm/s %wrqm w_await wareq-sz d/s dkB/s drqm/s %drqm d_await dareq-sz f/s f_await aqu-sz %util dm-0 400.60 14686.40 0.00 0.00 19.93 36.66 25.20 1197.60 0.00 0.00 27.44 47.52 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 8.67 91.32 dm-1 362.60 10583.20 0.00 0.00 18.58 29.19 30.40 1742.40 0.00 0.00 23.32 57.32 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 7.45 98.14 dm-10 6.40 64.00 0.00 0.00 0.03 10.00 37.60 150.40 0.00 0.00 0.06 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.02 dm-11 76.60 1939.20 0.00 0.00 10.61 25.32 35.20 1885.60 0.00 0.00 22.32 53.57 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.60 29.06 dm-12 93.00 4178.40 0.00 0.00 16.04 44.93 4.40 96.00 0.00 0.00 0.68 21.82 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.50 19.04 dm-13 376.80 12716.00 0.00 0.00 15.90 33.75 27.20 1444.80 0.00 0.00 23.76 53.12 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 6.64 97.76 dm-14 11.80 312.80 0.00 0.00 0.12 26.51 33.40 133.60 0.00 0.00 0.00 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.96 dm-15 7.00 430.40 0.00 0.00 0.17 61.49 46.20 184.80 0.00 0.00 0.06 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.94 dm-16 11.80 229.60 0.00 0.00 0.10 19.46 49.20 196.80 0.00 0.00 0.09 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 1.20 dm-17 4.00 39.20 0.00 0.00 0.10 9.80 29.60 118.40 0.00 0.00 0.00 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.72 dm-18 9.00 94.40 0.00 0.00 0.07 10.49 37.20 148.80 0.00 0.00 0.30 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 0.92 dm-19 7.40 70.40 0.00 0.00 0.11 9.51 56.80 227.20 0.00 0.00 0.02 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.04 dm-2 0.00 0.00 0.00 0.00 0.00 0.00 62.00 482.40 0.00 0.00 0.07 7.78 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.16 dm-20 12.40 2375.20 0.00 0.00 0.23 191.55 314.20 1524.80 0.00 0.00 0.03 4.85 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 7.90 dm-21 321.00 9468.00 0.00 0.00 7.86 29.50 31.20 1630.40 0.00 0.00 11.69 52.26 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 2.89 42.84 dm-3 12.40 125.60 0.00 0.00 0.08 10.13 37.00 148.00 0.00 0.00 0.02 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.84 dm-4 374.00 13150.40 0.00 0.00 18.31 35.16 23.60 1100.80 0.00 0.00 12.42 46.64 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 7.14 96.70 dm-5 8.00 84.00 0.00 0.00 0.10 10.50 38.00 152.00 0.00 0.00 0.03 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.00 dm-6 201.60 6619.20 0.00 0.00 11.56 32.83 3.40 108.80 0.00 0.00 2.29 32.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 2.34 58.24 dm-7 414.40 14476.00 0.00 0.00 21.99 34.93 10.80 235.20 0.00 0.00 5.19 21.78 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 9.17 98.00 dm-8 0.60 14.40 0.00 0.00 0.00 24.00 40.80 163.20 0.00 0.00 0.04 4.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.70 dm-9 478.00 17300.80 0.00 0.00 23.06 36.19 4.40 92.00 0.00 0.00 5.00 20.91 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 11.05 98.00 md127 5.00 64.80 0.00 0.00 0.04 12.96 133.80 2014.40 0.00 0.00 0.08 15.06 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 2.32 nvme0n1 33.60 822.40 0.00 0.00 0.09 24.48 252.40 3388.10 139.40 35.58 0.04 13.42 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.01 7.24 nvme1n1 62.60 3082.40 0.00 0.00 0.10 49.24 427.40 4272.90 177.20 29.31 0.03 10.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.02 13.70 sda 236.20 10613.60 126.20 34.82 15.19 44.93 11.80 1742.40 18.60 61.18 13.83 147.66 0.00 0.00 0.00 0.00 0.00 0.00 1.40 20.29 3.78 98.14 sdb 35.80 4166.40 56.80 61.34 10.03 116.38 4.20 96.00 0.20 4.55 0.81 22.86 0.00 0.00 0.00 0.00 0.00 0.00 0.80 0.25 0.36 19.04 sdc 118.20 6612.80 83.00 41.25 9.15 55.95 3.00 108.80 0.40 11.76 2.07 36.27 0.00 0.00 0.00 0.00 0.00 0.00 0.60 3.67 1.09 58.20 sdd 207.00 13152.00 166.60 44.59 13.92 63.54 15.80 1100.80 7.80 33.05 6.90 69.67 0.00 0.00 0.00 0.00 0.00 0.00 1.80 9.67 3.01 96.66 sde 55.80 1939.20 20.80 27.15 9.06 34.75 17.40 1885.60 17.80 50.57 7.51 108.37 0.00 0.00 0.00 0.00 0.00 0.00 1.80 1.44 0.64 29.04 sdf 221.40 12720.00 155.20 41.21 14.34 57.45 14.60 1444.80 12.60 46.32 10.51 98.96 0.00 0.00 0.00 0.00 0.00 0.00 1.40 11.71 3.35 97.70 sdg 234.60 14490.40 179.80 43.39 15.66 61.77 9.40 235.20 1.20 11.32 5.36 25.02 0.00 0.00 0.00 0.00 0.00 0.00 2.00 11.10 3.75 97.98 sdh 212.20 9470.40 109.20 33.98 4.05 44.63 19.40 1630.40 11.80 37.82 6.41 84.04 0.00 0.00 0.00 0.00 0.00 0.00 2.00 4.10 0.99 42.78 sdi 214.00 14684.00 186.60 46.58 14.46 68.62 10.00 1197.60 15.20 60.32 11.14 119.76 0.00 0.00 0.00 0.00 0.00 0.00 2.40 13.67 3.24 91.28 sdj 250.60 17296.80 227.00 47.53 18.76 69.02 3.60 92.00 0.80 18.18 5.72 25.56 0.00 0.00 0.00 0.00 0.00 0.00 0.80 16.25 4.73 98.00 `osd_max_scrubs` creates more reads per second (~120 -> ~220), but does not proportionately increase read throughput (as is expected on a spinning disk). Overall, this looks like scrub operations are seek-bound. Looking at `strace -fyp 640105 -e io_submit` output, for lines that involve my HDD `/dev/dm-12`: io_submit(0x7fd5f6668000, 2, [{aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd55baa3000, iov_len=4096}], aio_offset=1654046691328}, {aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f378000, iov_len=24576}], aio_offset=1654046666752}]) = 2 io_submit(0x7fd5f6668000, 2, [{aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f262000, iov_len=4096}], aio_offset=1934563819520}, {aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f28b000, iov_len=4096}], aio_offset=1934563823616}]) = 2 io_submit(0x7fd5f6668000, 2, [{aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f3c7000, iov_len=28672}], aio_offset=2871307956224}, {aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f2d2000, iov_len=4096}], aio_offset=2871307984896}]) = 2 io_submit(0x7fd5f6668000, 2, [{aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f34a000, iov_len=8192}], aio_offset=4056494669824}, {aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f35f000, iov_len=4096}], aio_offset=4056494678016}]) = 2 io_submit(0x7fd5f6668000, 2, [{aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f3e3000, iov_len=53248}], aio_offset=1233895051264}, {aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f0c6000, iov_len=4096}], aio_offset=1233895104512}]) = 2 io_submit(0x7fd5f6668000, 2, [{aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f1e2000, iov_len=4096}], aio_offset=2155792314368}, {aio_data=0, aio_lio_opcode=IOCB_CMD_PREADV, aio_fildes=30</dev/dm-12>, aio_buf=[{iov_base=0x7fd54f303000, iov_len=24576}], aio_offset=2155792289792}]) = 2 This looks like it's reading the HDD all over the place, with small reads (`aio_offset`, `iov_len`). Does Ceph scrubbing do a HDD seek for every object? I had imagined that scrubbing might be able to read linearly through the disk (perhaps skipping larger gaps). Is that wrong? Thanks!
I observed that on an otherwise idle cluster, scrubbing cannot fully utilise the speed of my HDDs.
Maybe the configured limit is set like this, because of that once (a part of) the scrubbing process is started it is not possible/easy to automatically scale down the performance to benefit client io.
`iostat` shows only 8-10 MB/s per disk, instead of the ~100 MB/s most HDDs can easily deliver.
100MB/s is sequential, your scrubbing is random. afaik everything is random.
Changing scrubbing settings does not help (see below).
I think you should be able to use the full performance of the disk when ceph tell osd.* injectargs '--osd_max_scrubs=X'. I never tried increasing the individual scrub speed, maybe such a setting is available ceph tell osd.* injectargs '--osd_recovery_sleep_hdd=0.100000' Best is to stick as much as possible to the defaults.
Hi Marc, thanks for your reply.
100MB/s is sequential, your scrubbing is random. afaik everything is random.
Is there any docs that explain this, any code, or other definitive answer? Also wouldn't it make sense that for scrubbing to be able to read the disk linearly, at least to some significant extent?
Changing scrubbing settings does not help (see below).
I think you should be able to use the full performance of the disk when ceph tell osd.* injectargs '--osd_max_scrubs=X'.
In my post I already showed that increasing `osd_max_scrubs` e.g. by 3x does not help. Also, what would be the logic how it could? If random IO is thrashing disk seeks, how could querying more concurrent disk seeks help?
ceph tell osd.* injectargs '--osd_recovery_sleep_hdd=0.100000'
There is no recovery going on in the cluster.
Hi Niklas,
100MB/s is sequential, your scrubbing is random. afaik everything is random.
Is there any docs that explain this, any code, or other definitive answer?
do a fio[1] test on a disk to see how it performs under certain conditions. Or look at atop during scrubbing, it will give you an impression how many % of your disk performance is used.
Also wouldn't it make sense that for scrubbing to be able to read the disk linearly, at least to some significant extent?
I would also think so, but I have no idea how this is implemented.
Changing scrubbing settings does not help (see below).
I think you should be able to use the full performance of the disk when ceph tell osd.* injectargs '--osd_max_scrubs=X'.
In my post I already showed that increasing `osd_max_scrubs` e.g. by 3x does not help.
Also, what would be the logic how it could?
I would argue. Because an individual scrub is not using all the disk resources. When you allow 2 scrub sessions on the same disk, it uses 2x the ios, which of course would be at the costs of available client io.
If random IO is thrashing disk seeks, how could querying more concurrent disk seeks help?
it is, but 1 single scrub session is not taking all of your disk io. None of the recovery procedures do, afaik. Because the cluster likes to serve client io first. The larger the cluster, the more often some part of the cluster is doing recovery.
ceph tell osd.* injectargs '--osd_recovery_sleep_hdd=0.100000'
There is no recovery going on in the cluster.
Yes I know, but this is a throttling factor, maybe something like this exists for scrubbing. The question you should ask yourself, why you want to change/investigate this? I like also to have a good performing cluster, but never looked at the scrubbing. Except turning it off before a reboot/update or so. [1] [global] ioengine=libaio #ioengine=posixaio invalidate=1 ramp_time=30 iodepth=1 runtime=180 time_based direct=1 filename=/dev/sdX #filename=/mnt/disk/fio-bench.img [write-4k-seq] stonewall bs=4k rw=write [randwrite-4k-seq] stonewall bs=4k rw=randwrite fsync=1 [read-4k-seq] stonewall bs=4k rw=read [randread-4k-seq] stonewall bs=4k rw=randread fsync=1 [rw-4k-seq] stonewall bs=4k rw=rw [randrw-4k-seq] stonewall bs=4k rw=randrw [randrw-4k-d4-seq] stonewall bs=4k rw=randrw iodepth=4 [randread-4k-d32-seq] stonewall bs=4k rw=randread iodepth=32 [randwrite-4k-d32-seq] stonewall bs=4k rw=randwrite iodepth=32 [write-128k-seq] stonewall bs=128k rw=write [randwrite-128k-seq] stonewall bs=128k rw=randwrite [read-128k-seq] stonewall bs=128k rw=read [randread-128k-seq] stonewall bs=128k rw=randread [rw-128k-seq] stonewall bs=128k rw=rw [randrw-128k-seq] stonewall bs=128k rw=randrw [write-1024k-seq] stonewall bs=1024k rw=write [randwrite-1024k-seq] stonewall bs=1024k rw=randwrite [read-1024k-seq] stonewall bs=1024k rw=read [randread-1024k-seq] stonewall bs=1024k rw=randread [rw-1024k-seq] stonewall bs=1024k rw=rw [randrw-1024k-seq] stonewall bs=1024k rw=randrw [write-4096k-seq] stonewall bs=4096k rw=write [write-4096k-d16-seq] stonewall bs=4M rw=write iodepth=16 [randwrite-4096k-seq] stonewall bs=4096k rw=randwrite [read-4096k-seq] stonewall bs=4096k rw=read [read-4096k-d16-seq] stonewall bs=4M rw=read iodepth=16 [randread-4096k-seq] stonewall bs=4096k rw=randread [rw-4096k-seq] stonewall bs=4096k rw=rw [randrw-4096k-seq] stonewall bs=4096k rw=randrw
The question you should ask yourself, why you want to change/investigate this?
Because if scrubbing takes 10x longer thrashing seeks, my scrubs never finish in time (the default is 1 week). I end with e.g.
267 pgs not deep-scrubbed in time
On a 38 TB cluster, if you scrub 8 MB/s on 10 disks (using only numbers already divided by replication factor), you need 55 days to scrub it once. That's 8x larger than the default scrub factor, so I'll get warnings and my risk of data degradation increases. Also, even if I set the default scrub interval to 8x larger, it my disks will still be thrashing seeks 100% of the time, affecting the cluster's throughput and latency performance. Niklas
On a 38 TB cluster, if you scrub 8 MB/s on 10 disks (using only numbers already divided by replication factor), you need 55 days to scrub it once. That's 8x larger than the default scrub factor [...] Also, even if I set the default scrub interval to 8x larger, it my disks will still be thrashing > seeks 100% of the time, affecting the cluster's throughput and latency performance.
Indeed! Every Ceph instance I have seen (not many) and almost every HPC storage system I have seen have this problem, and that's because they were never setup to have enough IOPS to support the maintenance load, never mind the maintenance load plus the user load (and as a rule not even the user load). There is a simple reason why this happens: when a large Ceph (etc. storage instance is initially setup, it is nearly empty, so it appears to perform well even if it was setup with inexpensive but slow/large HDDs, then it becomes fuller and therefore heavily congested but whoever set it up has already changed jobs or been promoted because of their initial success (or they invent excuses). A figure-of-merit that matters is IOPS-per-used-TB, and making it large enough to support concurrent maintenance (scrubbing, backfilling, rebalancing, backup) and user workloads. That is *expensive*, so in my experience very few storage instance buyers aim for that. The CERN IT people discovered long ago that quotes for storage workers always used very slow/large HDDs that performed very poorly if the specs were given as mere capacity, so they switched to requiring a different metric, 18MB/s transfer rate of *interleaved* read and write per TB of capacity, that is at least two parallel access streams per TB. https://www.sabi.co.uk/blog/13-two.html?131227#131227 "The issue with disk drives with multi-TB capacities" BTW I am not sure that a floor of 18MB/s of interleaved read and write per TB is high enough to support simultaneous maintenance and user loads for most Ceph instances, especially in HPC. I have seen HPC storage systems "designed" around 10TB and even 18TB HDDs, and the best that can be said about those HDDs is that they should be considered "tapes" with some random access ability.
The question you should ask yourself, why you want to change/investigate this?
Because if scrubbing takes 10x longer thrashing seeks, my scrubs never finish in time (the default is 1 week). I end with e.g.
267 pgs not deep-scrubbed in time
On a 38 TB cluster, if you scrub 8 MB/s on 10 disks (using only numbers already divided by replication factor), you need 55 days to scrub it once.
That's 8x larger than the default scrub factor, so I'll get warnings and my risk of data degradation increases.
Also, even if I set the default scrub interval to 8x larger, it my disks will still be thrashing seeks 100% of the time, affecting the cluster's throughput and latency performance.
Oh I get it. Interesting. I think if you will expand the cluster in the future with more disks you will spread the load have more iops, this will disappear. I am not sure if you will be able to fix this other than to increase the scrub interval. If you are sure it is nothing related to hardware. For you reference I have included how my disk io / performance looks like when I issue a deep-scrub. You can see it reads 2 disks here at ~70MB/s and the atop shows it is at 100% load. Nothing more you can do here. #ceph osd pool ls detail #ceph pg ls | grep '^53' #ceph osd tree #ceph pg deep-scrub 53.38 #dstat -D sdd,sde [@~]# dstat -d -D sdd,sde,sdj --dsk/sdd-----dsk/sde-----dsk/sdj-- read writ: read writ: read writ 2493k 177k:5086k 316k:5352k 422k 70M 0 : 89M 0 : 0 68k 78M 0 : 59M 0 : 0 0 68M 0 : 68M 0 : 0 28k 90M 4096B: 90M 80k:4096B 24k 76M 0 : 78M 0 : 0 12k 66M 0 : 64M 0 : 0 12k 70M 0 : 80M 0 :4096B 52k 77M 0 : 70M 0 : 0 0 atop: | DSK | sdd | busy 97% | read 1462 | write 4 | KiB/r 469 | KiB/w 5 | MBr/s 67.0 | MBw/s 0.0 | avq 1.01 | avio 6.59 ms | DSK | sde | busy 64% | read 1472 | write 4 | KiB/r 465 | KiB/w 6 | MBr/s 67.0 | MBw/s 0.0 | avq 1.01 | avio 4.32 ms | DSK | sdb | busy 1% | read 0 | write 82 | KiB/r 0 | KiB/w 9 | MBr/s 0.0 | MBw/s 0.1 | avq 1.30 | avio 1.29 ms |
The question you should ask yourself, why you want to change/investigate this?
Because if scrubbing takes 10x longer thrashing seeks, my scrubs never finish in time (the default is 1 week). I end with e.g.
267 pgs not deep-scrubbed in time
On a 38 TB cluster, if you scrub 8 MB/s on 10 disks (using only
numbers
already divided by replication factor), you need 55 days to scrub it once.
That's 8x larger than the default scrub factor, so I'll get warnings and my risk of data degradation increases.
Also, even if I set the default scrub interval to 8x larger, it my disks will still be thrashing seeks 100% of the time, affecting the cluster's throughput and latency performance.
Oh I get it. Interesting. I think if you will expand the cluster in the future with more disks you will spread the load have more iops, this will disappear. I am not sure if you will be able to fix this other than to increase the scrub interval. If you are sure it is nothing related to hardware.
For you reference I have included how my disk io / performance looks like when I issue a deep-scrub. You can see it reads 2 disks here at ~70MB/s and the atop shows it is at 100% load. Nothing more you can do here.
#ceph osd pool ls detail #ceph pg ls | grep '^53' #ceph osd tree #ceph pg deep-scrub 53.38 #dstat -D sdd,sde
[@~]# dstat -d -D sdd,sde,sdj --dsk/sdd-----dsk/sde-----dsk/sdj-- read writ: read writ: read writ 2493k 177k:5086k 316k:5352k 422k 70M 0 : 89M 0 : 0 68k 78M 0 : 59M 0 : 0 0 68M 0 : 68M 0 : 0 28k 90M 4096B: 90M 80k:4096B 24k 76M 0 : 78M 0 : 0 12k 66M 0 : 64M 0 : 0 12k 70M 0 : 80M 0 :4096B 52k 77M 0 : 70M 0 : 0 0
atop: | DSK | sdd | busy 97% | read 1462 | write 4 | KiB/r 469 | KiB/w 5 | MBr/s 67.0 | MBw/s 0.0 | avq 1.01 | avio 6.59 ms | DSK | sde | busy 64% | read 1472 | write 4 | KiB/r 465 | KiB/w 6 | MBr/s 67.0 | MBw/s 0.0 | avq 1.01 | avio 4.32 ms | DSK | sdb | busy 1% | read 0 | write 82 | KiB/r 0 | KiB/w 9 | MBr/s 0.0 | MBw/s 0.1 | avq 1.30 | avio 1.29 ms |
I did this on a pool with larger archived objects, when doing this on a filesystem with repo copies (rpm files) this performance is already dropping. DSK | sdh | busy 86% | read 1875 | write 26 | KiB/r 254 | KiB/w 4 | MBr/s 46.7 | MBw/s 0.0 | avq 1.59 | avio 4.50 ms DSK | sdd | busy 79% | read 1598 | write 63 | KiB/r 245 | KiB/w 16 | MBr/s 38.4 | MBw/s 0.1 | avq 1.89 | avio 4.77 ms DSK | sdf | busy 33% | read 1383 | write 139 | KiB/r 357 | KiB/w 7 | MBr/s 48.3 | MBw/s 0.1 | avq 1.14 | avio 2.20 ms
Hi Marc, thanks for your numbers, this seems to confirm the suspicions.
Oh I get it. Interesting. I think if you will expand the cluster in the future with more disks you will spread the load have more iops, this will disappear.
This one I'm not sure about: If I expand the cluster 2x, I'll also have 2x the data to scrub. So the ratio should be the same.
On a 38 TB cluster, if you scrub 8 MB/s on 10 disks (using only numbers already divided by replication factor), you need 55 days to scrub it once. That's 8x larger than the default scrub factor [...] Also, even if I set the default scrub interval to 8x larger, it my disks will still be thrashing seeks 100% of the time, affecting the cluster's throughput and latency performance.
Indeed! Every Ceph instance I have seen (not many) and almost every HPC storage system I have seen have this problem, and that's because they were never setup to have enough IOPS to support the maintenance load, never mind the maintenance load plus the user load (and as a rule not even the user load). There is a simple reason why this happens: when a large Ceph (etc. storage instance is initially setup, it is nearly empty, so it appears to perform well even if it was setup with inexpensive but slow/large HDDs, then it becomes fuller and therefore heavily congested but whoever set it up has already changed jobs or been promoted because of their initial success (or they invent excuses). A figure-of-merit that matters is IOPS-per-used-TB, and making it large enough to support concurrent maintenance (scrubbing, backfilling, rebalancing, backup) and user workloads. That is *expensive*, so in my experience very few storage instance buyers aim for that. The CERN IT people discovered long ago that quotes for storage workers always used very slow/large HDDs that performed very poorly if the specs were given as mere capacity, so they switched to requiring a different metric, 18MB/s transfer rate of *interleaved* read and write per TB of capacity, that is at least two parallel access streams per TB. https://www.sabi.co.uk/blog/13-two.html?131227#131227 "The issue with disk drives with multi-TB capacities" BTW I am not sure that a floor of 18MB/s of interleaved read and write per TB is high enough to support simultaneous maintenance and user loads for most Ceph instances, especially in HPC. I have seen HPC storage systems "designed" around 10TB and even 18TB HDDs, and the best that can be said about those HDDs is that they should be considered "tapes" with some random access ability.
Indeed! Every Ceph instance I have seen (not many) and almost every HPC storage system I have seen have this problem, and that's because they were never setup to have enough IOPS to support the maintenance load, never mind the maintenance load plus the user load (and as a rule not even the user load).
Yep, this is one of the false economies of spinners. The SNIA TCO calculator includes a performance factor for just this reason.
There is a simple reason why this happens: when a large Ceph (etc. storage instance is initially setup, it is nearly empty, so it appears to perform well even if it was setup with inexpensive but slow/large HDDs, then it becomes fuller and therefore heavily congested
Data fragments over time with organic growth, and the drive spends a larger fraction of time seeking. I’ve predicted then seen this even on a cluster whose hardware had been blessed by a certain professional services company (*ahem*).
but whoever set it up has already changed jobs or been promoted because of their initial success (or they invent excuses).
`xfs.mkfs -n size=65536` will haunt my nightmares until the end of my days. As well as an inadequate LFF HDD architecture I was not permitted to fix, *including the mons*. But I digress.
A figure-of-merit that matters is IOPS-per-used-TB, and making it large enough to support concurrent maintenance (scrubbing, backfilling, rebalancing, backup) and user workloads. That is *expensive*, so in my experience very few storage instance buyers aim for that.
^^^ This. Moreover, it’s all too common to try to band-aid this with expensive, fussy RoC HBAs with cache RAM and BBU/supercap. The money spent on those, and spent on jumping through their hoops, can easily debulk the HDD-SSD CapEx gap. Plus if your solution doesn’t do the job it needs to do, it is no bargain at any price. This correlates with IOPS/$, a metric in which HDDs are abysmal.
The CERN IT people discovered long ago that quotes for storage workers always used very slow/large HDDs that performed very poorly if the specs were given as mere capacity, so they switched to requiring a different metric, 18MB/s transfer rate of *interleaved* read and write per TB of capacity, that is at least two parallel access streams per TB.
At least one major SSD manufacturer attends specifically to reads under write pressure.
https://www.sabi.co.uk/blog/13-two.html?131227#131227 "The issue with disk drives with multi-TB capacities"
BTW I am not sure that a floor of 18MB/s of interleaved read and write per TB is high enough to support simultaneous maintenance and user loads for most Ceph instances, especially in HPC.
I have seen HPC storage systems "designed" around 10TB and even 18TB HDDs, and the best that can be said about those HDDs is that they should be considered "tapes" with some random access ability.
Yes! This harks back to DECtape https://www.vt100.net/timeline/1964.html which was literally this, people even used it at a filesystem. Some years ago I had Brian Kernighan sign one “Wow I haven’t seen one of these in YEARS!” — aad
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
On a 38 TB cluster, if you scrub 8 MB/s on 10 disks (using only numbers already divided by replication factor), you need 55 days to scrub it once. That's 8x larger than the default scrub factor [...] Also, even if I set the default scrub interval to 8x larger, it my disks will still be thrashing seeks 100% of the time, affecting the cluster's throughput and latency performance.
Indeed! Every Ceph instance I have seen (not many) and almost every HPC storage system I have seen have this problem, and that's because they were never setup to have enough IOPS to support the maintenance load, never mind the maintenance load plus the user load (and as a rule not even the user load). There is a simple reason why this happens: when a large Ceph (etc. storage instance is initially setup, it is nearly empty, so it appears to perform well even if it was setup with inexpensive but slow/large HDDs, then it becomes fuller and therefore heavily congested but whoever set it up has already changed jobs or been promoted because of their initial success (or they invent excuses). A figure-of-merit that matters is IOPS-per-used-TB, and making it large enough to support concurrent maintenance (scrubbing, backfilling, rebalancing, backup) and user workloads. That is *expensive*, so in my experience very few storage instance buyers aim for that. The CERN IT people discovered long ago that quotes for storage workers always used very slow/large HDDs that performed very poorly if the specs were given as mere capacity, so they switched to requiring a different metric, 18MB/s transfer rate of *interleaved* read and write per TB of capacity, that is at least two parallel access streams per TB. https://www.sabi.co.uk/blog/13-two.html?131227#131227 "The issue with disk drives with multi-TB capacities" BTW I am not sure that a floor of 18MB/s of interleaved read and write per TB is high enough to support simultaneous maintenance and user loads for most Ceph instances, especially in HPC. I have seen HPC storage systems "designed" around 10TB and even 18TB HDDs, and the best that can be said about those HDDs is that they should be considered "tapes" with some random access ability.
Hi, I asked a similar question about increasing scrub throughput some time ago and couldn't get a fully satisfying answer: https://lists.ceph.io/hyperkitty/list/ceph-users@ceph.io/thread/NHOHZLVQ3CKM... My observation is that much fewer (deep) scrubs are scheduled than could be executed. Some people wrote scripts to do scrub scheduling in a more efficient way (by last-scrub time stamp), but I don't want to go this route (yet). Unfortunately, the thread above does not contain the full conversation, I think it forked into a second one with the same or a similar title. About performance calculations, along the lines of
they were never setup to have enough IOPS to support the maintenance load, never mind the maintenance load plus the user load
initially setup, it is nearly empty, so it appears to perform well even if it was setup with inexpensive but slow/large HDDs, then it becomes fuller and therefore heavily congested
There is a bit more to that. HDDs have the unfortunate property that sector reads/writes are not independent of which sector is read/written to. An empty drive will serve IO from the beginning of the disk when everything is fast. As drives fill up, they start using slower and slower regions. This performance degradation is in addition to the effects of longer seek paths and fragmentation. Here I'm talking only about enterprise data centre drives with proper sustained performance profiles, not cheap stuff that falls apart once you go serious. Unfortunately, ceph adds on top of that the lack of tail merging support, which makes small objects extra expensive. Still, ceph was written for HDDs and actually performs well if IO calculations are done properly. For example, 8TB vs. 18TB drives. 8TB drives start with about 150MB/s bandwidth at the fast part and slow down to 80-100MB/s when you reach the end. 18TB drives are not just 8TB drives with denser packing, they actually have more platters. That means, they start out at 250MB/s and reach something like 100-130MB/s towards the end. Its more than double the capacity, but not more than double the throughput. IOP/s are roughly the same, so IOP/s per TB go down a lot with capacity. When is this fine and when is it problematic. Its fine if you have large objects that are never modified. Then ceph will usually reach sequential read/write performance and scrubbing will be done within a week (with less than 10% utilisation, which is good). The other extreme is many small objects, in which case your observed performance/throughput can be terrible and scrubbing might never end. For being able to make reasonable estimates, you need to know real-life object size distributions and if full object writes are effectively sequential (meaning you have large bluestore alloc sizes in general, look at the bluestore performance counters, it will indicate how many large and how many small writes you have). We have a fairly mixed size distribution with, unfortunately, quite a percentage of small objects on our ceph fs. We do have 18T drives, which are about 30% utilised. Scrubbing still finishes within less than 2 weeks even with the outliers due to "not ideal" scrub scheduling (thread above). I'm willing to accept up to 4 weeks tail time, which will probably give me 50-60% utilisation before things go below acceptable. In essence, the 18T average performance drives are something like 10T pretty good performance drives compared with the usual 8T drives. You just have to let go of 100% capacity utilisation. The limit is what comes first, capacity- or IOP/s saturation. Once admin workload cannot complete in time, that's it, the disks are full and one needs to expand. We have about 900 HDDs in our cluster and I maintain this large number mostly for performance reasons. I don't think I will ever see more than 50% utilisation before we change deployment or add drives. Looking at our data in more detail, most of it is ice cold. Therefore, in the long run we plan to go for tiered OSDs (bcache/dm-cache) with sufficient total SSD capacity to hold about 2 times all hot data. Then, maybe, we can fill big drives a bit more. I was looking into large capacity SSDs and, I'm afraid, when going to the >=18TB SSD section they either have bad and often worse performance than spinners, or are massively expensive. With performance here I mean bandwidth. Large SSDs can have a sustained bandwith of 30MB/s. They will still do about 500-1000IOP/s per TB, but large file transfer or backfill will become a pain. I looked at models with reasonable bandwidth and asked if I could get a price. The answer was that one such disk costs more than an entire of our standard storage servers. Clearly not our league. A better solution is to combine the best of both worlds and have a more intelligent software that can differentiate between hot and cold data and may be able to adapt to workloads.
the best that can be said about those HDDs is that they should be considered "tapes" with some random access ability
Which is good if that is all you need. But true, a lot of people already forget that using an 8+3 EC profile on a pool will divide the aggregated IOP/s budget by 11. After this, divide by 2 and you have a number to tell your users/boss. They are either happy or give you more money. Our users also think in terms of price/TB only. I simply incorporate performance into the calculation and come up with price per *usable* TB. Raw capacity includes admin overhead (which includes IOP/s), which can easily be 50% in total plus the replication overhead. Just let go of 100% capacity utilisation and you will have a well working cluster. I let go of 50% utilisation. That's when I start requesting material and it works really well. Still much cheaper than an all-flash install with higher utilisation. To the all-flash enthusiasts. Yes, we have all-flash pools and I do enjoy their performance. Still, the price. There are people who say platters are outdated and SSDs are competitive. Well, my google-fu is maybe not good enough, so here we go. If you show me where I can get SSDs with the specs below, I will go all-flash. Until then, sorry, cost economy is still a thing. Specs A: - capacity: 18TB+ - sustained 1M block-size sequential read/write (iodepth=1): 15MB/s per TB - sustained 4K random 50/50 read-write (iodepth=1): 100 - data written per day for 5 years: 1TB (yes, this *is* very low yet sufficient) - interface: SATA/SAS, 2.5" or 3.5" - price: <=350$ (for 18TB) Specs B: - capacity: 18TB+ - sustained 1M block-size sequential read/write (iodepth=1): 25MB/s per TB - sustained 4K random 50/50 read-write (iodepth=1): 1000 - data written per day for 5 years: 1TB (yes, this *is* very low yet sufficient) - interface: SATA/SAS, 2.5" or 3.5" - price: <=700$ (for 18TB) Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14 ________________________________________ From: Peter Grandi <pg@ceph.list.sabi.co.uk> Sent: Thursday, April 27, 2023 11:55 AM To: list fs Ceph Subject: [ceph-users] Re: Deep-scrub much slower than HDD speed
On a 38 TB cluster, if you scrub 8 MB/s on 10 disks (using only numbers already divided by replication factor), you need 55 days to scrub it once. That's 8x larger than the default scrub factor [...] Also, even if I set the default scrub interval to 8x larger, it my disks will still be thrashing seeks 100% of the time, affecting the cluster's throughput and latency performance.
Indeed! Every Ceph instance I have seen (not many) and almost every HPC storage system I have seen have this problem, and that's because they were never setup to have enough IOPS to support the maintenance load, never mind the maintenance load plus the user load (and as a rule not even the user load). There is a simple reason why this happens: when a large Ceph (etc. storage instance is initially setup, it is nearly empty, so it appears to perform well even if it was setup with inexpensive but slow/large HDDs, then it becomes fuller and therefore heavily congested but whoever set it up has already changed jobs or been promoted because of their initial success (or they invent excuses). A figure-of-merit that matters is IOPS-per-used-TB, and making it large enough to support concurrent maintenance (scrubbing, backfilling, rebalancing, backup) and user workloads. That is *expensive*, so in my experience very few storage instance buyers aim for that. The CERN IT people discovered long ago that quotes for storage workers always used very slow/large HDDs that performed very poorly if the specs were given as mere capacity, so they switched to requiring a different metric, 18MB/s transfer rate of *interleaved* read and write per TB of capacity, that is at least two parallel access streams per TB. https://www.sabi.co.uk/blog/13-two.html?131227#131227 "The issue with disk drives with multi-TB capacities" BTW I am not sure that a floor of 18MB/s of interleaved read and write per TB is high enough to support simultaneous maintenance and user loads for most Ceph instances, especially in HPC. I have seen HPC storage systems "designed" around 10TB and even 18TB HDDs, and the best that can be said about those HDDs is that they should be considered "tapes" with some random access ability. _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Den ons 26 apr. 2023 kl 21:20 skrev Niklas Hambüchen <mail@nh2.me>:
100MB/s is sequential, your scrubbing is random. afaik everything is random.
Is there any docs that explain this, any code, or other definitive answer? Also wouldn't it make sense that for scrubbing to be able to read the disk linearly, at least to some significant extent?
Scrubs only read data that does exist in ceph as it exists, not every sector of the drive, written or not. This is why many small objects make it look like MB/s is "low", it reads the objects and not just dumb cylinder reads. It is not the same as hw raid boxes doing "patrol reads" or what they call it where they have no idea of what they are reading, just seeing that the drives don't report errors. This is more "pretend you are a ceph client with low priority reading all the data from this PG from start to end". -- May the most significant bit of your life be positive.
Hi all,
Scrubs only read data that does exist in ceph as it exists, not every sector of the drive, written or not.
Thanks, this does explain it. I just discovered: ZFS had this problem in the past: * https://utcc.utoronto.ca/~cks/space/blog/solaris/ZFSNonlinearScrubs?showcomm... OpenZFS solved it in 2017, using two-phase scrubs: * https://github.com/openzfs/zfs/issues/3625 * https://github.com/openzfs/zfs/commit/d4a72f23863382bdf6d0ae33196f5b5decbc48... Perhaps Ceph can use the same approach; I filed https://tracker.ceph.com/issues/59584 for it.
Den fre 28 apr. 2023 kl 14:51 skrev Niklas Hambüchen <niklas@benaco.com>:
Hi all,
Scrubs only read data that does exist in ceph as it exists, not every sector of the drive, written or not.
Thanks, this does explain it.
I just discovered:
ZFS had this problem in the past:
* https://utcc.utoronto.ca/~cks/space/blog/solaris/ZFSNonlinearScrubs?showcomm...
That one talks about resilvering, which is not the same as neither ZFS scrubs nor ceph scrubs. Resilvering is copying/moving/rewriting data, which scrubs are not. -- May the most significant bit of your life be positive.
participants (7)
-
Anthony D'Atri
-
Frank Schilder
-
Janne Johansson
-
Marc
-
Niklas Hambüchen
-
Niklas Hambüchen
-
Peter Grandi