March 2021 - ceph-users - lists.ceph.io

by Ramanathan S

Hi all, I just had created a ceph cluster to use cephfs. When i create the a ceph fs pool i get the filesystem below error. # ceph osd pool create cephfs_data 128 pool 'cephfs_data' created # ceph osd pool create cephfs_metadata 128 pool 'cephfs_metadata' created # ceph fs new cephfs cephfs_metadata cephfs_data new fs with metadata pool 6 and data pool 5 # ceph -s cluster: id: 1c27def45-f0f9-494d-sfke-eb4323432fd health: HEALTH_ERR 1 filesystem is offline 1 filesystem is online with fewer MDS than max_mds services: mon: 2 daemons, quorum ceph-mon01,ceph-mon02 mgr: ceph-adm01(active) mds: cephfs-0/0/1 up osd: 12 osds: 12 up, 12 in data: pools: 2 pools, 256 pgs objects: 0 objects, 0 B usage: 12 GiB used, 588 GiB / 600 GiB avail pgs: 256 active+clean but when i check the max_mds for the ceph fs it says 1 # ceph fs get cephfs | grep max_mds max_mds 1 Let anyone know what am i missing here? Any inputs is much appreciated. Regards, Ram Ceph-explorer..

3 weeks, 3 days

3
3
0 0

Re: NoSuchKey on key that is visible in s3 list/radosgw bk

by Eric Ivancich

I have some questions for those who’ve experienced this issue. 1. It seems like those reporting this issue are seeing it strictly after upgrading to Octopus. From what version did each of these sites upgrade to Octopus? From Nautilus? Mimic? Luminous? 2. Does anyone have any lifecycle rules on a bucket experiencing this issue? If so, please describe. 3. Is anyone making copies of the affected objects (to same or to a different bucket) prior to seeing the issue? And if they are making copies, does the destination bucket have lifecycle rules? And if they are making copies, are those copies ever being removed? 4. Is anyone experiencing this issue willing to run their RGWs with 'debug_ms=1'? That would allow us to see a request from an RGW to either remove a tail object or decrement its reference counter (and when its counter reaches 0 it will be deleted). Thanks, Eric > On Nov 12, 2020, at 4:54 PM, huxiaoyu(a)horebdata.cn wrote: > > Looks like this is a very dangerous bug for data safety. Hope the bug would be quickly identified and fixed. > > best regards, > > Samuel > > > > huxiaoyu(a)horebdata.cn <mailto:huxiaoyu@horebdata.cn> > > From: Janek Bevendorff > Date: 2020-11-12 18:17 > To: huxiaoyu(a)horebdata.cn <mailto:huxiaoyu@horebdata.cn>; EDH - Manuel Rios; Rafael Lopez > CC: Robin H. Johnson; ceph-users > Subject: Re: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk > I have never seen this on Luminous. I recently upgraded to Octopus and the issue started occurring only few weeks later. > > On 12/11/2020 16:37, huxiaoyu(a)horebdata.cn wrote: > which Ceph versions are affected by this RGW bug/issues? Luminous, Mimic, Octupos, or the latest? > > any idea? > > samuel > > > > huxiaoyu(a)horebdata.cn > > From: EDH - Manuel Rios > Date: 2020-11-12 14:27 > To: Janek Bevendorff; Rafael Lopez > CC: Robin H. Johnson; ceph-users > Subject: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk > This same error caused us to wipe a full cluster of 300TB... will be related to some rados index/database bug not to s3. > > As Janek exposed is a mayor issue, because the error silent happend and you can only detect it with S3, when you're going to delete/purge a S3 bucket. Dropping NoSuchKey. Error is not related to S3 logic .. > > Hope this time dev's can take enought time to find and resolve the issue. Error happens with low ec profiles, even with replica x3 in some cases. > > Regards > > > > -----Mensaje original----- > De: Janek Bevendorff <janek.bevendorff(a)uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>> > Enviado el: jueves, 12 de noviembre de 2020 14:06 > Para: Rafael Lopez <rafael.lopez(a)monash.edu <mailto:rafael.lopez@monash.edu>> > CC: Robin H. Johnson <robbat2(a)gentoo.org <mailto:robbat2@gentoo.org>>; ceph-users <ceph-users(a)ceph.io <mailto:ceph-users@ceph.io>> > Asunto: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk > > Here is a bug report concerning (probably) this exact issue: > https://tracker.ceph.com/issues/47866 <https://tracker.ceph.com/issues/47866> > > I left a comment describing the situation and my (limited) experiences with it. > > > On 11/11/2020 10:04, Janek Bevendorff wrote: >> >> Yeah, that seems to be it. There are 239 objects prefixed >> .8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh in my dump. However, there are none >> of the multiparts from the other file to be found and the head object >> is 0 bytes. >> >> I checked another multipart object with an end pointer of 11. >> Surprisingly, it had way more than 11 parts (39 to be precise) named >> .1, .1_1 .1_2, .1_3, etc. Not sure how Ceph identifies those, but I >> could find them in the dump at least. >> >> I have no idea why the objects disappeared. I ran a Spark job over all >> buckets, read 1 byte of every object and recorded errors. Of the 78 >> buckets, two are missing objects. One bucket is missing one object, >> the other 15. So, luckily, the incidence is still quite low, but the >> problem seems to be expanding slowly. >> >> >> On 10/11/2020 23:46, Rafael Lopez wrote: >>> Hi Janek, >>> >>> What you said sounds right - an S3 single part obj won't have an S3 >>> multipart string as part of the prefix. S3 multipart string looks >>> like "2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme". >>> >>> From memory, single part S3 objects that don't fit in a single rados >>> object are assigned a random prefix that has nothing to do with >>> the object name, and the rados tail/data objects (not the head >>> object) have that prefix. >>> As per your working example, the prefix for that would be >>> '.8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh'. So there would be (239) "shadow" >>> objects with names containing that prefix, and if you add up the >>> sizes it should be the size of your S3 object. >>> >>> You should look at working and non working examples of both single >>> and multipart S3 objects, as they are probably all a bit different >>> when you look in rados. >>> >>> I agree it is a serious issue, because once objects are no longer in >>> rados, they cannot be recovered. If it was a case that there was a >>> link broken or rados objects renamed, then we could work to >>> recover...but as far as I can tell, it looks like stuff is just >>> vanishing from rados. The only explanation I can think of is some >>> (rgw or rados) background process is incorrectly doing something with >>> these objects (eg. renaming/deleting). I had thought perhaps it was a >>> bug with the rgw garbage collector..but that is pure speculation. >>> >>> Once you can articulate the problem, I'd recommend logging a bug >>> tracker upstream. >>> >>> >>> On Wed, 11 Nov 2020 at 06:33, Janek Bevendorff >>> <janek.bevendorff(a)uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de> >>> <mailto:janek.bevendorff@uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>>> wrote: >>> >>> Here's something else I noticed: when I stat objects that work >>> via radosgw-admin, the stat info contains a "begin_iter" JSON >>> object with RADOS key info like this >>> >>> >>> "key": { >>> "name": >>> "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", >>> "instance": "", >>> "ns": "" >>> } >>> >>> >>> and then "end_iter" with key info like this: >>> >>> >>> "key": { >>> "name": >>> ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", >>> "instance": "", >>> "ns": "shadow" >>> } >>> >>> However, when I check the broken 0-byte object, the "begin_iter" >>> and "end_iter" keys look like this: >>> >>> >>> "key": { >>> "name": >>> "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", >>> "instance": "", >>> "ns": "multipart" >>> } >>> >>> [...] >>> >>> >>> "key": { >>> "name": >>> "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", >>> "instance": "", >>> "ns": "multipart" >>> } >>> >>> So, it's the full name plus a suffix and the namespace is >>> multipart, not shadow (or empty). This in itself may just be an >>> artefact of whether the object was uploaded in one go or as a >>> multipart object, but the second difference is that I cannot find >>> any of the multipart objects in my pool's object name dump. I >>> can, however, find the shadow RADOS object of the intact S3 object. >>> >>> >>> >>> >>> -- >>> *Rafael Lopez* >>> Devops Systems Engineer >>> Monash University eResearch Centre >>> >>> T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> >>> E: rafael.lopez(a)monash.edu <mailto:rafael.lopez@monash.edu> >>> > _______________________________________________ > ceph-users mailing list -- ceph-users(a)ceph.io > To unsubscribe send an email to ceph-users-leave(a)ceph.io > _______________________________________________ > ceph-users mailing list -- ceph-users(a)ceph.io > To unsubscribe send an email to ceph-users-leave(a)ceph.io > _______________________________________________ > ceph-users mailing list -- ceph-users(a)ceph.io > To unsubscribe send an email to ceph-users-leave(a)ceph.io

3 weeks, 3 days

6
22
0 0

Small RGW objects and RADOS 64KB minimun size

by Loïc Dachary

Bonjour, Reading Karan's blog post about benchmarking the insertion of billions objects to Ceph via S3 / RGW[0] from last year, it reads: > we decided to lower bluestore_min_alloc_size_hdd to 18KB and re-test. As represented in chart-5, the object creation rate found to be notably reduced after lowering the bluestore_min_alloc_size_hdd parameter from 64KB (default) to 18KB. As such, for objects larger than the bluestore_min_alloc_size_hdd , the default values seems to be optimal, smaller objects further require more investigation if you intended to reduce bluestore_min_alloc_size_hdd parameter. There also is a mail thread dated 2018 on this topic as well, with the same conclusion although using RADOS directly and not RGW[3]. I read the RGW data layout page in the documentation[1] and concluded that by default every object inserted with S3 / RGW will indeed use at least 64kb. A pull request from last year[2] seems to confirm it and also suggests modifying bluestore_min_alloc_size_hdd has adverse side effects. That being said, I'm curious to know if people developed strategies to cope with this overhead. Someone mentioned packing objects together client side to make them larger. But maybe there are simpler ways to do the same? Cheers [0] https://www.redhat.com/en/blog/scaling-ceph-billion-objects-and-beyond [1] https://docs.ceph.com/en/latest/radosgw/layout/ [2] https://github.com/ceph/ceph/pull/32809 [3] https://www.spinics.net/lists/ceph-users/msg45755.html -- Loïc Dachary, Artisan Logiciel Libre

11 months

4
6
0 0

cephadm cluster move /var/lib/docker to separate device fails

by Karsten Nielsen

Hi, I have setup a ceph cluster with cephadm with docker backend. I want to move /var/lib/docker to a separate device to get better performance and less load on the OS device. I tried that by stopping docker copy the content of /var/lib/docker to the new device and mount the new device to /var/lib/docker. The other containers started as expected and continues to work and run as expected. But the ceph containers seems to be broken. I am not able to get them back in working state. I have tried to remove the host with `ceph orch host rm itcnchn-bb4067` and readd it but no effect. The strange thing is that 2 of 4 containers comes up as expected. ceph orch ps itcnchn-bb4067 NAME HOST STATUS REFRESHED AGE VERSION IMAGE NAME IMAGE ID CONTAINER ID crash.itcnchn-bb4067 itcnchn-bb4067 running (18h) 10m ago 4w 15.2.7 docker.io/ceph/ceph:v15 2bc420ddb175 2af28c4571cf mds.cephfs.itcnchn-bb4067.qzoshl itcnchn-bb4067 error 10m ago 4w <unknown> docker.io/ceph/ceph:v15 <unknown> <unknown> mon.itcnchn-bb4067 itcnchn-bb4067 error 10m ago 18h <unknown> docker.io/ceph/ceph:v15 <unknown> <unknown> rgw.ikea.dc9-1.itcnchn-bb4067.gtqedc itcnchn-bb4067 running (18h) 10m ago 4w 15.2.7 docker.io/ceph/ceph:v15 2bc420ddb175 00d000aec32b Docker logs from the active manager does not say much about what is wrong debug 2021-01-05T09:57:52.537+0000 7fdb69691700 0 log_channel(cephadm) log [INF] : Reconfiguring mds.cephfs.itcnchn-bb4067.qzoshl (unknown last config time)... debug 2021-01-05T09:57:52.541+0000 7fdb69691700 0 log_channel(cephadm) log [INF] : Reconfiguring daemon mds.cephfs.itcnchn-bb4067.qzoshl on itcnchn-bb4067 debug 2021-01-05T09:57:52.973+0000 7fdb64e88700 0 log_channel(cluster) log [DBG] : pgmap v347: 241 pgs: 241 active+clean; 18 GiB data, 50 GiB used, 52 TiB / 52 TiB avail; 18 KiB/s rd, 78 KiB/s wr, 24 op/s debug 2021-01-05T09:57:53.085+0000 7fdb69691700 0 log_channel(cephadm) log [INF] : Reconfiguring mon.itcnchn-bb4067 (unknown last config time)... debug 2021-01-05T09:57:53.085+0000 7fdb69691700 0 log_channel(cephadm) log [INF] : Reconfiguring daemon mon.itcnchn-bb4067 on itcnchn-bb4067 debug 2021-01-05T09:57:53.625+0000 7fdb69691700 0 log_channel(cephadm) log [INF] : Reconfiguring rgw.ikea.dc9-1.itcnchn-bb4067.gtqedc (unknown last config time)... debug 2021-01-05T09:57:53.629+0000 7fdb69691700 0 log_channel(cephadm) log [INF] : Reconfiguring daemon rgw.ikea.dc9-1.itcnchn-bb4067.gtqedc on itcnchn-bb4067 debug 2021-01-05T09:57:54.141+0000 7fdb69691700 0 log_channel(cephadm) log [INF] : Reconfiguring crash.itcnchn-bb4067 (unknown last config time)... debug 2021-01-05T09:57:54.141+0000 7fdb69691700 0 log_channel(cephadm) log [INF] : Reconfiguring daemon crash.itcnchn-bb4067 on itcnchn-bb4067 - Karsten

1 year, 1 month

2
1
0 0

kernel client osdc ops stuck and mds slow reqs

by Dan van der Ster

Hi all, We are quite regularly (a couple times per week) seeing: HEALTH_WARN 1 clients failing to respond to capability release; 1 MDSs report slow requests MDS_CLIENT_LATE_RELEASE 1 clients failing to respond to capability release mdshpc-be143(mds.0): Client hpc-be028.cern.ch: failing to respond to capability release client_id: 52919162 MDS_SLOW_REQUEST 1 MDSs report slow requests mdshpc-be143(mds.0): 1 slow requests are blocked > 30 secs Which is being caused by osdc ops stuck in a kernel client, e.g.: 10:57:18 root hpc-be028 /root → cat /sys/kernel/debug/ceph/4da6fd06-b069-49af-901f-c9513baabdbd.client52919162/osdc REQUESTS 9 homeless 0 46559317 osd243 3.ee6ffcdb 3.cdb [243,501,92]/243 [243,501,92]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f09a01.00000057 0x400014 1 read 46559322 osd243 3.ee6ffcdb 3.cdb [243,501,92]/243 [243,501,92]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f09a01.00000057 0x400014 1 read 46559323 osd243 3.969cc573 3.573 [243,330,226]/243 [243,330,226]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f09a56.00000056 0x400014 1 read 46559341 osd243 3.969cc573 3.573 [243,330,226]/243 [243,330,226]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f09a56.00000056 0x400014 1 read 46559342 osd243 3.969cc573 3.573 [243,330,226]/243 [243,330,226]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f09a56.00000056 0x400014 1 read 46559345 osd243 3.969cc573 3.573 [243,330,226]/243 [243,330,226]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f09a56.00000056 0x400014 1 read 46559621 osd243 3.6313e8ef 3.8ef [243,330,521]/243 [243,330,521]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f09a45.0000007a 0x400014 1 read 46559629 osd243 3.b280c852 3.852 [243,113,539]/243 [243,113,539]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f09a3a.0000007f 0x400014 1 read 46559928 osd243 3.1ee7bab4 3.ab4 [243,332,94]/243 [243,332,94]/243 e678697 fsvolumens_355f485c-6319-4ffe-acd6-94a07f2a14b4/10003f099ff.0000073f 0x400024 1 write LINGER REQUESTS BACKOFFS We can unblock those requests by doing `ceph osd down osd.243` (or restarting osd.243). This is ceph v14.2.6 and the client kernel is el7 3.10.0-957.27.2.el7.x86_64. Are there a better way to debug this? Best Regards, Dan

1 year, 2 months

4
12
0 0

cephfs forward scrubbing docs

by Dan van der Ster

Hi, Today while debugging something we had a few questions that might lead to improving the cephfs forward scrub docs: https://docs.ceph.com/en/latest/cephfs/scrub/ tldr: 1. Should we document which sorts of issues that the forward scrub is able to fix? 2. Can we make it more visible (in docs) that scrubbing is not supported with multi-mds? 3. Isn't the new `ceph -s` scrub task status misleading with multi-mds? Details here: 1) We found a CephFS directory with a number of zero sized files: # ls -l ... -rw-r--r-- 1 1001890000 1001890000 0 Nov 3 11:58 upload_fc501199e3e7abe6b574101cf34aeefb.png -rw-r--r-- 1 1001890000 1001890000 0 Nov 3 12:23 upload_fce4f55348185fefa0abdd8d11095ba8.gif -rw-r--r-- 1 1001890000 1001890000 0 Nov 3 11:54 upload_fd95b8358851f0dac22fb775046a6163.png ... The user claims that those files were non-zero sized last week. The sequence of zero sized files includes *all* files written between Nov 2 and 9. The user claims that his client was running out of memory, but this is now fixed. So I suspect that his ceph client (kernel 3.10.0-1127.19.1.el7.x86_64) was not behaving well. Anyway, I noticed that even though the dentries list 0 bytes, the underlying rados objects have data, and the data looks good. E.g. # rados get -p cephfs_data 200212e68b5.00000000 --namespace=xxx 200212e68b5.00000000 # file 200212e68b5.00000000 200212e68b5.00000000: PNG image data, 960 x 815, 8-bit/color RGBA, non-interlaced So I managed to recover the files doing something like this (using an input file mapping inode to filename) [see PS 0]. But I'm wondering if a forward scrub is able to fix this sort of problem directly? Should we document which sorts of issues that the forward scrub is able to fix? I anyway tried to scrub it, which led to: # ceph tell mds.cephflax-mds-xxx scrub start /volumes/_nogroup/xxx recursive repair Scrub is not currently supported for multiple active MDS. Please reduce max_mds to 1 and then scrub. So ... 2) Shouldn't we update the doc to mention loud and clear that scrub is not currently supported for multiple active MDS? 3) I was somehow surprised by this, because I had thought that the new `ceph -s` multi-mds scrub status implied that multi-mds scrubbing was now working: task status: scrub status: mds.x: idle mds.y: idle mds.z: idle Is it worth reporting this task status for cephfs if we can't even scrub them? Thanks!! Dan [0] mkdir -p recovered while read -r a b; do for i in {0..9} do echo "rados stat --cluster=flax --pool=cephfs_data --namespace=xxx" $(printf "%x" $a).0000000$i "&&" "rados get --cluster=flax --pool=cephfs_data --namespace=xxx" $(printf "%x" $a).0000000$i $(printf "%x" $a).0000000$i done echo cat $(printf "%x" $a).* ">" $(printf "%x" $a) echo mv $(printf "%x" $a) recovered/$b done < inones_fnames.txt

2 years, 10 months

2
2
0 0

does ceph rgw has any option to limit bandwidth

by Zhenshi Zhou

Hi, Is there any option of rados gateway that limit bandwidth?

2 years, 10 months

5
6
0 0

Recommendations on problem with PG

by Gabriel Medve

Hi, We have a problem with a PG that was inconsistent, currently the PG in our cluster have 3 copies. It was not possible for us to repair this pg with "ceph pg repair" (This PG is in osd 14,1,2) so we deleted some of the copies of osd 14 with the following command. ceph-objectstore-tool --data-path /var/lib/ceph/osd.14/ --pgid 22.f --op remove --force This caused an automatic attempt to create the missing copy entering the backfilling state, but when doing this it crashed osd 1 and 2 and threw the IOPS to 0, freezing the cluster. Is there any way to remove this entire pg or try to recreate the missing copy or ignore it completely? It causes instability in the cluster. Thank you, I await comments -- Untitled Document ------------------------------------------------------------------------ Gabriel I. Medve

2 years, 11 months

2
2
0 0

ceph-ansible in Pacific and beyond?

by Matthew Vernon

Hi, I caught up with Sage's talk on what to expect in Pacific ( https://www.youtube.com/watch?v=PVtn53MbxTc ) and there was no mention of ceph-ansible at all. Is it going to continue to be supported? We use it (and uncontainerised packages) for all our clusters, so I'd be a bit alarmed if it was going to go away... Regards, Matthew -- The Wellcome Sanger Institute is operated by Genome Research Limited, a charity registered in England with number 1021457 and a company registered in England with number 2742969, whose registered office is 215 Euston Road, London, NW1 2BE.

2 years, 11 months

17
25
0 0

RGW: Multiple Site does not sync olds data

by 特木勒

Hi all: ceph version: 15.2.7 (88e41c6c49beb18add4fdb6b4326ca466d931db8) I have a strange question, I just create a multiple site for Ceph cluster. But I notice the old data of source cluster is not synced. Only new data will be synced into second zone cluster. Is there anything I need to do to enable full sync for bucket or this is a bug? Thanks

2 years, 11 months

3
17
0 0

2024

2023

2022

2021

2020

2019

ceph-users March 2021