Re: NoSuchKey on key that is visible in s3 list/radosgw bk
We are having the exact same problem (also Octopus). The object is listed by s3cmd, but trying to download it results in a 404 error. radosgw-admin object stat shows that the object still exists. Any further ideas how I can restore access to this object? (Sorry if this is a duplicate, but it seems like the mailing list hasn't accepted my original mail).
Mariusz Gronczewski wrote:
Dnia 2020-07-27, o godz. 21:31:33 "Robin H. Johnson" <robbat2@gentoo.org <mailto:robbat2@gentoo.org>> napisał(a):
On Mon, Jul 27, 2020 at 08:02:23PM +0200, Mariusz Gronczewski wrote:
Hi, I've got a problem on Octopus (15.2.3, debian packages) install, bucket S3 index shows a file: s3cmd ls s3://upvid/255/38355 --recursive 2020-07-27 17:48 50584342 s3://upvid/255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4 radosgw-admin bi list also shows it { "type": "plain", "idx": "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", "entry": { "name": "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", "instance": "", "ver": { "pool": 11, "epoch": 853842 }, "locator": "", "exists": "true", "meta": { "category": 1, "size": 50584342, "mtime": "2020-07-27T17:48:27.203008Z", "etag": "2b31cc8ce8b1fb92a5f65034f2d12581-7", "storage_class": "", "owner": "filmweb-app", "owner_display_name": "filmweb app user", "content_type": "", "accounted_size": 50584342, "user_data": "", "appendable": "false" }, "tag": "_3ubjaztglHXfZr05wZCFCPzebQf-ZFP", "flags": 0, "pending_map": [], "versioned_epoch": 0 } }, but trying to download it via curl (I've set permissions to public0
only gets me Does the RADOS object for this still exist?
try: radosgw-admin object stat --bucket ... --object '255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4'
If that doesn't return, then the backing object is gone, and you have a stale index entry that can be cleaned up in most cases with check bucket. For cases where that doesn't fix it, my recommended way to fix it is write a new 0-byte object to the same name, then delete it.
it does exist:
{ "name": "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", "size": 50584342, "policy": { "acl": {...}, "owner": {...} }, "etag": "2b31cc8ce8b1fb92a5f65034f2d12581-7", "tag": "_3ubjaztglHXfZr05wZCFCPzebQf-ZFP", "manifest": { "objs": [], "obj_size": 50584342, "explicit_objs": "false", "head_size": 0, "max_head_size": 0, "prefix": "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2", "rules": [ { "key": 0, "val": { "start_part_num": 1, "start_ofs": 0, "part_size": 8388608, "stripe_max_size": 4194304, "override_prefix": "" } }, { "key": 50331648, "val": { "start_part_num": 7, "start_ofs": 50331648, "part_size": 252694, "stripe_max_size": 4194304, "override_prefix": "" } } ], "tail_instance": "", "tail_placement": { "bucket": { "name": "upvid", "marker": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", "explicit_placement": { "data_pool": "", "data_extra_pool": "", "index_pool": "" } }, "placement_rule": "default-placement" }, "begin_iter": { "part_ofs": 0, "stripe_ofs": 0, "ofs": 0, "stripe_size": 4194304, "cur_part_id": 1, "cur_stripe": 0, "cur_override_prefix": "", "location": { "placement_rule": "default-placement", "obj": { "bucket": { "name": "upvid", "marker": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", "explicit_placement": { "data_pool": "", "data_extra_pool": "", "index_pool": "" } }, "key": { "name": "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2.1", "instance": "", "ns": "multipart" } }, "raw_obj": { "pool": "", "oid": "", "loc": "" }, "is_raw": false } }, "end_iter": { "part_ofs": 50584342, "stripe_ofs": 50584342, "ofs": 50584342, "stripe_size": 252694, "cur_part_id": 8, "cur_stripe": 0, "cur_override_prefix": "", "location": { "placement_rule": "default-placement", "obj": { "bucket": { "name": "upvid", "marker": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", "explicit_placement": { "data_pool": "", "data_extra_pool": "", "index_pool": "" } }, "key": { "name": "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2.8", "instance": "", "ns": "multipart" } }, "raw_obj": { "pool": "", "oid": "", "loc": "" }, "is_raw": false } } }, "attrs": { "user.rgw.pg_ver": "/u0005u", "user.rgw.source_zone": "�I}�", "user.rgw.tail_tag": "88d4f221-0da5-444d-81a8-517771278350.658638.5538809", "user.rgw.x-amz-acl": "private", "user.rgw.x-amz-content-sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", "user.rgw.x-amz-date": "20200619T194517Z" } }
-- Mariusz Gronczewski, Administrator
Efigence S. A. ul. Wołoska 9a, 02-583 Warszawa T: [+48] 22 380 13 13 NOC: [+48] 22 380 10 20 E: admin@efigence.com <mailto:admin@efigence.com> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
Hi Mariusz, all We have seen this issue as well, on redhat ceph 4 (I have an unresolved case open). In our case, `radosgw-admin stat` is not a sufficient check to guarantee that there are rados objects. You have to do a `rados stat` to know that. In your case, the object is ~48M in size, appears to also use S3 multipart. This means, when uploaded, S3 will slice it up into parts based on what S3 multipart size you use (5M default, i think 8M here). After that, rados further slices any incoming (multipart size objects) into rados object objects of 4Mib size (default). The end result is you have a bunch of rados objects labelled with the 'prefix' from the `radosgw-admin stat` you ran, as well as a head object (named the same as the S3 object you uploaded) that contains the metadata so rgw knows how to put the S3 object back together. In our case, the head object is there but the other rados pieces that hold the actual data seem to be gone, so `radosgw-admin stat` returns fine, but we get NoSuchKey when trying to download. Try `rados -p {rgw buckets pool} stat 255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4`, it will show you the rados stat of the head object, which will be much smaller than the S3 object. To check if you actually have all rados objects for this 48M S3 object, try searching for parts of the prefix or the whole prefix on a list of all rados objects in buckets pool. FYI, the `rados ls` will list every rados object in the bucket, so it may be very large and take a long time if you have many objects. rados -p {rgw buckets pool} ls > {tmpfile} grep '2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2' {tmpfile} grep 'juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4' {tmpfile} The first grep is actually the S3 multipart ID string added to the prefix by rgw. Rafael On Tue, 10 Nov 2020 at 01:04, Janek Bevendorff < janek.bevendorff@uni-weimar.de> wrote:
We are having the exact same problem (also Octopus). The object is listed by s3cmd, but trying to download it results in a 404 error. radosgw-admin object stat shows that the object still exists. Any further ideas how I can restore access to this object?
(Sorry if this is a duplicate, but it seems like the mailing list hasn't accepted my original mail).
Mariusz Gronczewski wrote:
Dnia 2020-07-27, o godz. 21:31:33 "Robin H. Johnson" <robbat2@gentoo.org <mailto:robbat2@gentoo.org>> napisał(a):
On Mon, Jul 27, 2020 at 08:02:23PM +0200, Mariusz Gronczewski wrote:
Hi, I've got a problem on Octopus (15.2.3, debian packages) install, bucket S3 index shows a file: s3cmd ls s3://upvid/255/38355 --recursive 2020-07-27 17:48 50584342
s3://upvid/255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4
radosgw-admin bi list also shows it { "type": "plain", "idx":
"255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4",
"entry": { "name":
"255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4",
"instance": "", "ver": { "pool": 11, "epoch": 853842 }, "locator": "", "exists": "true", "meta": { "category": 1, "size": 50584342, "mtime": "2020-07-27T17:48:27.203008Z", "etag": "2b31cc8ce8b1fb92a5f65034f2d12581-7", "storage_class": "", "owner": "filmweb-app", "owner_display_name": "filmweb app user", "content_type": "", "accounted_size": 50584342, "user_data": "", "appendable": "false" }, "tag": "_3ubjaztglHXfZr05wZCFCPzebQf-ZFP", "flags": 0, "pending_map": [], "versioned_epoch": 0 } }, but trying to download it via curl (I've set permissions to public0
only gets me Does the RADOS object for this still exist?
try: radosgw-admin object stat --bucket ... --object '255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4'
If that doesn't return, then the backing object is gone, and you have a stale index entry that can be cleaned up in most cases with check bucket. For cases where that doesn't fix it, my recommended way to fix it is write a new 0-byte object to the same name, then delete it.
it does exist:
{ "name": "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", "size": 50584342, "policy": { "acl": {...}, "owner": {...} }, "etag": "2b31cc8ce8b1fb92a5f65034f2d12581-7", "tag": "_3ubjaztglHXfZr05wZCFCPzebQf-ZFP", "manifest": { "objs": [], "obj_size": 50584342, "explicit_objs": "false", "head_size": 0, "max_head_size": 0, "prefix":
"255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2",
"rules": [ { "key": 0, "val": { "start_part_num": 1, "start_ofs": 0, "part_size": 8388608, "stripe_max_size": 4194304, "override_prefix": "" } }, { "key": 50331648, "val": { "start_part_num": 7, "start_ofs": 50331648, "part_size": 252694, "stripe_max_size": 4194304, "override_prefix": "" } } ], "tail_instance": "", "tail_placement": { "bucket": { "name": "upvid", "marker": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", "explicit_placement": { "data_pool": "", "data_extra_pool": "", "index_pool": "" } }, "placement_rule": "default-placement" }, "begin_iter": { "part_ofs": 0, "stripe_ofs": 0, "ofs": 0, "stripe_size": 4194304, "cur_part_id": 1, "cur_stripe": 0, "cur_override_prefix": "", "location": { "placement_rule": "default-placement", "obj": { "bucket": { "name": "upvid", "marker": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", "explicit_placement": { "data_pool": "", "data_extra_pool": "", "index_pool": "" } }, "key": { "name":
"255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2.1",
"instance": "", "ns": "multipart" } }, "raw_obj": { "pool": "", "oid": "", "loc": "" }, "is_raw": false } }, "end_iter": { "part_ofs": 50584342, "stripe_ofs": 50584342, "ofs": 50584342, "stripe_size": 252694, "cur_part_id": 8, "cur_stripe": 0, "cur_override_prefix": "", "location": { "placement_rule": "default-placement", "obj": { "bucket": { "name": "upvid", "marker": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", "explicit_placement": { "data_pool": "", "data_extra_pool": "", "index_pool": "" } }, "key": { "name":
"255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2.8",
"instance": "", "ns": "multipart" } }, "raw_obj": { "pool": "", "oid": "", "loc": "" }, "is_raw": false } } }, "attrs": { "user.rgw.pg_ver": "/u0005u", "user.rgw.source_zone": "�I}�", "user.rgw.tail_tag": "88d4f221-0da5-444d-81a8-517771278350.658638.5538809", "user.rgw.x-amz-acl": "private", "user.rgw.x-amz-content-sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", "user.rgw.x-amz-date": "20200619T194517Z" } }
-- Mariusz Gronczewski, Administrator
Efigence S. A. ul. Wołoska 9a, 02-583 Warszawa T: [+48] 22 380 13 13 NOC: [+48] 22 380 10 20 E: admin@efigence.com <mailto:admin@efigence.com> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto: ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre T: +61 3 9905 9118 E: rafael.lopez@monash.edu
Thanks for the reply. This issue seems to be VERY serious. New objects are disappearing every day. This is a silent, creeping data loss. I couldn't find the object with rados stat, but I am now listing all the objects and will grep the dump to see if there is anything left. Janek On 09/11/2020 23:31, Rafael Lopez wrote:
Hi Mariusz, all
We have seen this issue as well, on redhat ceph 4 (I have an unresolved case open). In our case, `radosgw-admin stat` is not a sufficient check to guarantee that there are rados objects. You have to do a `rados stat` to know that.
In your case, the object is ~48M in size, appears to also use S3 multipart. This means, when uploaded, S3 will slice it up into parts based on what S3 multipart size you use (5M default, i think 8M here). After that, rados further slices any incoming (multipart size objects) into rados object objects of 4Mib size (default).
The end result is you have a bunch of rados objects labelled with the 'prefix' from the `radosgw-admin stat` you ran, as well as a head object (named the same as the S3 object you uploaded) that contains the metadata so rgw knows how to put the S3 object back together. In our case, the head object is there but the other rados pieces that hold the actual data seem to be gone, so `radosgw-admin stat` returns fine, but we get NoSuchKey when trying to download.
Try `rados -p {rgw buckets pool} stat 255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4`, it will show you the rados stat of the head object, which will be much smaller than the S3 object.
To check if you actually have all rados objects for this 48M S3 object, try searching for parts of the prefix or the whole prefix on a list of all rados objects in buckets pool. FYI, the `rados ls` will list every rados object in the bucket, so it may be very large and take a long time if you have many objects.
rados -p {rgw buckets pool} ls > {tmpfile} grep '2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2' {tmpfile} grep 'juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4' {tmpfile}
The first grep is actually the S3 multipart ID string added to the prefix by rgw.
Rafael
On Tue, 10 Nov 2020 at 01:04, Janek Bevendorff <janek.bevendorff@uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>> wrote:
We are having the exact same problem (also Octopus). The object is listed by s3cmd, but trying to download it results in a 404 error. radosgw-admin object stat shows that the object still exists. Any further ideas how I can restore access to this object?
(Sorry if this is a duplicate, but it seems like the mailing list hasn't accepted my original mail).
> Mariusz Gronczewski wrote: > > >> Dnia 2020-07-27, o godz. 21:31:33 >> "Robin H. Johnson" <robbat2@gentoo.org <mailto:robbat2@gentoo.org> >> <mailto:robbat2@gentoo.org <mailto:robbat2@gentoo.org>>> napisał(a): >> >> >>> On Mon, Jul 27, 2020 at 08:02:23PM +0200, Mariusz Gronczewski wrote: >>> >>>> Hi, >>>> I've got a problem on Octopus (15.2.3, debian packages) install, >>>> bucket S3 index shows a file: >>>> s3cmd ls s3://upvid/255/38355 --recursive >>>> 2020-07-27 17:48 50584342 >>>> s3://upvid/255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4 >>>> radosgw-admin bi list also shows it >>>> { >>>> "type": "plain", >>>> "idx": >>>> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", >>>> "entry": { "name": >>>> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", >>>> "instance": "", "ver": { >>>> "pool": 11, >>>> "epoch": 853842 >>>> }, >>>> "locator": "", >>>> "exists": "true", >>>> "meta": { >>>> "category": 1, >>>> "size": 50584342, >>>> "mtime": "2020-07-27T17:48:27.203008Z", >>>> "etag": "2b31cc8ce8b1fb92a5f65034f2d12581-7", >>>> "storage_class": "", >>>> "owner": "filmweb-app", >>>> "owner_display_name": "filmweb app user", >>>> "content_type": "", >>>> "accounted_size": 50584342, >>>> "user_data": "", >>>> "appendable": "false" >>>> }, >>>> "tag": "_3ubjaztglHXfZr05wZCFCPzebQf-ZFP", >>>> "flags": 0, >>>> "pending_map": [], >>>> "versioned_epoch": 0 >>>> } >>>> }, >>>> but trying to download it via curl (I've set permissions to public0 >>>> >>> only gets me >>> Does the RADOS object for this still exist? >>> >>> try: >>> radosgw-admin object stat --bucket ... --object >>> '255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4' >>> >>> If that doesn't return, then the backing object is gone, and you have >>> a stale index entry that can be cleaned up in most cases with check >>> bucket. >>> For cases where that doesn't fix it, my recommended way to fix it is >>> write a new 0-byte object to the same name, then delete it. >>> >> >> >> >> it does exist: >> >> { >> "name": >> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", >> "size": 50584342, "policy": { >> "acl": {...}, >> "owner": {...} >> }, >> "etag": "2b31cc8ce8b1fb92a5f65034f2d12581-7", >> "tag": "_3ubjaztglHXfZr05wZCFCPzebQf-ZFP", >> "manifest": { >> "objs": [], >> "obj_size": 50584342, >> "explicit_objs": "false", >> "head_size": 0, >> "max_head_size": 0, >> "prefix": >> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2", >> "rules": [ { >> "key": 0, >> "val": { >> "start_part_num": 1, >> "start_ofs": 0, >> "part_size": 8388608, >> "stripe_max_size": 4194304, >> "override_prefix": "" >> } >> }, >> { >> "key": 50331648, >> "val": { >> "start_part_num": 7, >> "start_ofs": 50331648, >> "part_size": 252694, >> "stripe_max_size": 4194304, >> "override_prefix": "" >> } >> } >> ], >> "tail_instance": "", >> "tail_placement": { >> "bucket": { >> "name": "upvid", >> "marker": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", >> "explicit_placement": { >> "data_pool": "", >> "data_extra_pool": "", >> "index_pool": "" >> } >> }, >> "placement_rule": "default-placement" >> }, >> "begin_iter": { >> "part_ofs": 0, >> "stripe_ofs": 0, >> "ofs": 0, >> "stripe_size": 4194304, >> "cur_part_id": 1, >> "cur_stripe": 0, >> "cur_override_prefix": "", >> "location": { >> "placement_rule": "default-placement", >> "obj": { >> "bucket": { >> "name": "upvid", >> "marker": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", >> "explicit_placement": { >> "data_pool": "", >> "data_extra_pool": "", >> "index_pool": "" >> } >> }, >> "key": { >> "name": >> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2.1", >> "instance": "", "ns": "multipart" >> } >> }, >> "raw_obj": { >> "pool": "", >> "oid": "", >> "loc": "" >> }, >> "is_raw": false >> } >> }, >> "end_iter": { >> "part_ofs": 50584342, >> "stripe_ofs": 50584342, >> "ofs": 50584342, >> "stripe_size": 252694, >> "cur_part_id": 8, >> "cur_stripe": 0, >> "cur_override_prefix": "", >> "location": { >> "placement_rule": "default-placement", >> "obj": { >> "bucket": { >> "name": "upvid", >> "marker": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", >> "explicit_placement": { >> "data_pool": "", >> "data_extra_pool": "", >> "index_pool": "" >> } >> }, >> "key": { >> "name": >> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2.8", >> "instance": "", "ns": "multipart" >> } >> }, >> "raw_obj": { >> "pool": "", >> "oid": "", >> "loc": "" >> }, >> "is_raw": false >> } >> } >> }, >> "attrs": { >> "user.rgw.pg_ver": "/u0005u", >> "user.rgw.source_zone": "�I}�", >> "user.rgw.tail_tag": >> "88d4f221-0da5-444d-81a8-517771278350.658638.5538809", >> "user.rgw.x-amz-acl": "private", "user.rgw.x-amz-content-sha256": >> "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", >> "user.rgw.x-amz-date": "20200619T194517Z" } >> } >> >> >> >> >> >> -- >> Mariusz Gronczewski, Administrator >> >> Efigence S. A. >> ul. Wołoska 9a, 02-583 Warszawa >> T: [+48] 22 380 13 13 >> NOC: [+48] 22 380 10 20 >> E: admin@efigence.com <mailto:admin@efigence.com> <mailto:admin@efigence.com <mailto:admin@efigence.com>> >> _______________________________________________ >> ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> <mailto:ceph-users@ceph.io <mailto:ceph-users@ceph.io>> >> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io> >> <mailto:ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre
T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> E: rafael.lopez@monash.edu <mailto:rafael.lopez@monash.edu>
I found some of the data in the rados ls dump. We host some WARCs from the Internet Archive and one affected WARC still has its warc.os.cdx.gz file intact, while the actual warc.gz is gone. A rados stat revealed WIDE-20110903143858-01166.warc.os.cdx.gz mtime 2019-07-14T17:48:39.000000+0200, size 1060428 for the cdx.gz file, but WIDE-20110903143858-01166.warc.gz mtime 2019-07-14T17:04:49.000000+0200, size 0 for the warc.gz. I couldn't find any of the suffixed multipart objects listed in radosgw-admin stat. WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19: (2) No such file or directory On 10/11/2020 10:14, Janek Bevendorff wrote:
Thanks for the reply. This issue seems to be VERY serious. New objects are disappearing every day. This is a silent, creeping data loss.
I couldn't find the object with rados stat, but I am now listing all the objects and will grep the dump to see if there is anything left.
Janek
On 09/11/2020 23:31, Rafael Lopez wrote:
Hi Mariusz, all
We have seen this issue as well, on redhat ceph 4 (I have an unresolved case open). In our case, `radosgw-admin stat` is not a sufficient check to guarantee that there are rados objects. You have to do a `rados stat` to know that.
In your case, the object is ~48M in size, appears to also use S3 multipart. This means, when uploaded, S3 will slice it up into parts based on what S3 multipart size you use (5M default, i think 8M here). After that, rados further slices any incoming (multipart size objects) into rados object objects of 4Mib size (default).
The end result is you have a bunch of rados objects labelled with the 'prefix' from the `radosgw-admin stat` you ran, as well as a head object (named the same as the S3 object you uploaded) that contains the metadata so rgw knows how to put the S3 object back together. In our case, the head object is there but the other rados pieces that hold the actual data seem to be gone, so `radosgw-admin stat` returns fine, but we get NoSuchKey when trying to download.
Try `rados -p {rgw buckets pool} stat 255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4`, it will show you the rados stat of the head object, which will be much smaller than the S3 object.
To check if you actually have all rados objects for this 48M S3 object, try searching for parts of the prefix or the whole prefix on a list of all rados objects in buckets pool. FYI, the `rados ls` will list every rados object in the bucket, so it may be very large and take a long time if you have many objects.
rados -p {rgw buckets pool} ls > {tmpfile} grep '2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2' {tmpfile} grep 'juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4' {tmpfile}
The first grep is actually the S3 multipart ID string added to the prefix by rgw.
Rafael
On Tue, 10 Nov 2020 at 01:04, Janek Bevendorff <janek.bevendorff@uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>> wrote:
We are having the exact same problem (also Octopus). The object is listed by s3cmd, but trying to download it results in a 404 error. radosgw-admin object stat shows that the object still exists. Any further ideas how I can restore access to this object?
(Sorry if this is a duplicate, but it seems like the mailing list hasn't accepted my original mail).
> Mariusz Gronczewski wrote: > > >> Dnia 2020-07-27, o godz. 21:31:33 >> "Robin H. Johnson" <robbat2@gentoo.org <mailto:robbat2@gentoo.org> >> <mailto:robbat2@gentoo.org <mailto:robbat2@gentoo.org>>> napisał(a): >> >> >>> On Mon, Jul 27, 2020 at 08:02:23PM +0200, Mariusz Gronczewski wrote: >>> >>>> Hi, >>>> I've got a problem on Octopus (15.2.3, debian packages) install, >>>> bucket S3 index shows a file: >>>> s3cmd ls s3://upvid/255/38355 --recursive >>>> 2020-07-27 17:48 50584342 >>>> s3://upvid/255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4 >>>> radosgw-admin bi list also shows it >>>> { >>>> "type": "plain", >>>> "idx": >>>> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", >>>> "entry": { "name": >>>> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", >>>> "instance": "", "ver": { >>>> "pool": 11, >>>> "epoch": 853842 >>>> }, >>>> "locator": "", >>>> "exists": "true", >>>> "meta": { >>>> "category": 1, >>>> "size": 50584342, >>>> "mtime": "2020-07-27T17:48:27.203008Z", >>>> "etag": "2b31cc8ce8b1fb92a5f65034f2d12581-7", >>>> "storage_class": "", >>>> "owner": "filmweb-app", >>>> "owner_display_name": "filmweb app user", >>>> "content_type": "", >>>> "accounted_size": 50584342, >>>> "user_data": "", >>>> "appendable": "false" >>>> }, >>>> "tag": "_3ubjaztglHXfZr05wZCFCPzebQf-ZFP", >>>> "flags": 0, >>>> "pending_map": [], >>>> "versioned_epoch": 0 >>>> } >>>> }, >>>> but trying to download it via curl (I've set permissions to public0 >>>> >>> only gets me >>> Does the RADOS object for this still exist? >>> >>> try: >>> radosgw-admin object stat --bucket ... --object >>> '255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4' >>> >>> If that doesn't return, then the backing object is gone, and you have >>> a stale index entry that can be cleaned up in most cases with check >>> bucket. >>> For cases where that doesn't fix it, my recommended way to fix it is >>> write a new 0-byte object to the same name, then delete it. >>> >> >> >> >> it does exist: >> >> { >> "name": >> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4", >> "size": 50584342, "policy": { >> "acl": {...}, >> "owner": {...} >> }, >> "etag": "2b31cc8ce8b1fb92a5f65034f2d12581-7", >> "tag": "_3ubjaztglHXfZr05wZCFCPzebQf-ZFP", >> "manifest": { >> "objs": [], >> "obj_size": 50584342, >> "explicit_objs": "false", >> "head_size": 0, >> "max_head_size": 0, >> "prefix": >> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2", >> "rules": [ { >> "key": 0, >> "val": { >> "start_part_num": 1, >> "start_ofs": 0, >> "part_size": 8388608, >> "stripe_max_size": 4194304, >> "override_prefix": "" >> } >> }, >> { >> "key": 50331648, >> "val": { >> "start_part_num": 7, >> "start_ofs": 50331648, >> "part_size": 252694, >> "stripe_max_size": 4194304, >> "override_prefix": "" >> } >> } >> ], >> "tail_instance": "", >> "tail_placement": { >> "bucket": { >> "name": "upvid", >> "marker": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", >> "explicit_placement": { >> "data_pool": "", >> "data_extra_pool": "", >> "index_pool": "" >> } >> }, >> "placement_rule": "default-placement" >> }, >> "begin_iter": { >> "part_ofs": 0, >> "stripe_ofs": 0, >> "ofs": 0, >> "stripe_size": 4194304, >> "cur_part_id": 1, >> "cur_stripe": 0, >> "cur_override_prefix": "", >> "location": { >> "placement_rule": "default-placement", >> "obj": { >> "bucket": { >> "name": "upvid", >> "marker": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", >> "explicit_placement": { >> "data_pool": "", >> "data_extra_pool": "", >> "index_pool": "" >> } >> }, >> "key": { >> "name": >> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2.1", >> "instance": "", "ns": "multipart" >> } >> }, >> "raw_obj": { >> "pool": "", >> "oid": "", >> "loc": "" >> }, >> "is_raw": false >> } >> }, >> "end_iter": { >> "part_ofs": 50584342, >> "stripe_ofs": 50584342, >> "ofs": 50584342, >> "stripe_size": 252694, >> "cur_part_id": 8, >> "cur_stripe": 0, >> "cur_override_prefix": "", >> "location": { >> "placement_rule": "default-placement", >> "obj": { >> "bucket": { >> "name": "upvid", >> "marker": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "bucket_id": >> "88d4f221-0da5-444d-81a8-517771278350.665933.2", "tenant": "", >> "explicit_placement": { >> "data_pool": "", >> "data_extra_pool": "", >> "index_pool": "" >> } >> }, >> "key": { >> "name": >> "255/38355/juz_nie_zyjesz_sezon_2___oficjalny_zwiastun___netflix_mp4.2~NTy88SkDkXR9ifSrrRcw5WPDxqN3PO2.8", >> "instance": "", "ns": "multipart" >> } >> }, >> "raw_obj": { >> "pool": "", >> "oid": "", >> "loc": "" >> }, >> "is_raw": false >> } >> } >> }, >> "attrs": { >> "user.rgw.pg_ver": "/u0005u", >> "user.rgw.source_zone": "�I}�", >> "user.rgw.tail_tag": >> "88d4f221-0da5-444d-81a8-517771278350.658638.5538809", >> "user.rgw.x-amz-acl": "private", "user.rgw.x-amz-content-sha256": >> "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", >> "user.rgw.x-amz-date": "20200619T194517Z" } >> } >> >> >> >> >> >> -- >> Mariusz Gronczewski, Administrator >> >> Efigence S. A. >> ul. Wołoska 9a, 02-583 Warszawa >> T: [+48] 22 380 13 13 >> NOC: [+48] 22 380 10 20 >> E: admin@efigence.com <mailto:admin@efigence.com> <mailto:admin@efigence.com <mailto:admin@efigence.com>> >> _______________________________________________ >> ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> <mailto:ceph-users@ceph.io <mailto:ceph-users@ceph.io>> >> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io> >> <mailto:ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io>
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre
T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> E: rafael.lopez@monash.edu <mailto:rafael.lopez@monash.edu>
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Here's something else I noticed: when I stat objects that work via radosgw-admin, the stat info contains a "begin_iter" JSON object with RADOS key info like this "key": { "name": "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", "instance": "", "ns": "" } and then "end_iter" with key info like this: "key": { "name": ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", "instance": "", "ns": "shadow" } However, when I check the broken 0-byte object, the "begin_iter" and "end_iter" keys look like this: "key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", "instance": "", "ns": "multipart" } [...] "key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", "instance": "", "ns": "multipart" } So, it's the full name plus a suffix and the namespace is multipart, not shadow (or empty). This in itself may just be an artefact of whether the object was uploaded in one go or as a multipart object, but the second difference is that I cannot find any of the multipart objects in my pool's object name dump. I can, however, find the shadow RADOS object of the intact S3 object.
Hi Janek, What you said sounds right - an S3 single part obj won't have an S3 multipart string as part of the prefix. S3 multipart string looks like "2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme". From memory, single part S3 objects that don't fit in a single rados object are assigned a random prefix that has nothing to do with the object name, and the rados tail/data objects (not the head object) have that prefix. As per your working example, the prefix for that would be '.8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh'. So there would be (239) "shadow" objects with names containing that prefix, and if you add up the sizes it should be the size of your S3 object. You should look at working and non working examples of both single and multipart S3 objects, as they are probably all a bit different when you look in rados. I agree it is a serious issue, because once objects are no longer in rados, they cannot be recovered. If it was a case that there was a link broken or rados objects renamed, then we could work to recover...but as far as I can tell, it looks like stuff is just vanishing from rados. The only explanation I can think of is some (rgw or rados) background process is incorrectly doing something with these objects (eg. renaming/deleting). I had thought perhaps it was a bug with the rgw garbage collector..but that is pure speculation. Once you can articulate the problem, I'd recommend logging a bug tracker upstream. On Wed, 11 Nov 2020 at 06:33, Janek Bevendorff < janek.bevendorff@uni-weimar.de> wrote:
Here's something else I noticed: when I stat objects that work via radosgw-admin, the stat info contains a "begin_iter" JSON object with RADOS key info like this
"key": { "name": "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", "instance": "", "ns": "" }
and then "end_iter" with key info like this:
"key": { "name": ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", "instance": "", "ns": "shadow" }
However, when I check the broken 0-byte object, the "begin_iter" and "end_iter" keys look like this:
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", "instance": "", "ns": "multipart" }
[...]
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", "instance": "", "ns": "multipart" }
So, it's the full name plus a suffix and the namespace is multipart, not shadow (or empty). This in itself may just be an artefact of whether the object was uploaded in one go or as a multipart object, but the second difference is that I cannot find any of the multipart objects in my pool's object name dump. I can, however, find the shadow RADOS object of the intact S3 object.
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre T: +61 3 9905 9118 E: rafael.lopez@monash.edu
Yeah, that seems to be it. There are 239 objects prefixed .8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh in my dump. However, there are none of the multiparts from the other file to be found and the head object is 0 bytes. I checked another multipart object with an end pointer of 11. Surprisingly, it had way more than 11 parts (39 to be precise) named .1, .1_1 .1_2, .1_3, etc. Not sure how Ceph identifies those, but I could find them in the dump at least. I have no idea why the objects disappeared. I ran a Spark job over all buckets, read 1 byte of every object and recorded errors. Of the 78 buckets, two are missing objects. One bucket is missing one object, the other 15. So, luckily, the incidence is still quite low, but the problem seems to be expanding slowly. On 10/11/2020 23:46, Rafael Lopez wrote:
Hi Janek,
What you said sounds right - an S3 single part obj won't have an S3 multipart string as part of the prefix. S3 multipart string looks like "2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme".
From memory, single part S3 objects that don't fit in a single rados object are assigned a random prefix that has nothing to do with the object name, and the rados tail/data objects (not the head object) have that prefix. As per your working example, the prefix for that would be '.8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh'. So there would be (239) "shadow" objects with names containing that prefix, and if you add up the sizes it should be the size of your S3 object.
You should look at working and non working examples of both single and multipart S3 objects, as they are probably all a bit different when you look in rados.
I agree it is a serious issue, because once objects are no longer in rados, they cannot be recovered. If it was a case that there was a link broken or rados objects renamed, then we could work to recover...but as far as I can tell, it looks like stuff is just vanishing from rados. The only explanation I can think of is some (rgw or rados) background process is incorrectly doing something with these objects (eg. renaming/deleting). I had thought perhaps it was a bug with the rgw garbage collector..but that is pure speculation.
Once you can articulate the problem, I'd recommend logging a bug tracker upstream.
On Wed, 11 Nov 2020 at 06:33, Janek Bevendorff <janek.bevendorff@uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>> wrote:
Here's something else I noticed: when I stat objects that work via radosgw-admin, the stat info contains a "begin_iter" JSON object with RADOS key info like this
"key": { "name": "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", "instance": "", "ns": "" }
and then "end_iter" with key info like this:
"key": { "name": ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", "instance": "", "ns": "shadow" }
However, when I check the broken 0-byte object, the "begin_iter" and "end_iter" keys look like this:
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", "instance": "", "ns": "multipart" }
[...]
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", "instance": "", "ns": "multipart" }
So, it's the full name plus a suffix and the namespace is multipart, not shadow (or empty). This in itself may just be an artefact of whether the object was uploaded in one go or as a multipart object, but the second difference is that I cannot find any of the multipart objects in my pool's object name dump. I can, however, find the shadow RADOS object of the intact S3 object.
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre
T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> E: rafael.lopez@monash.edu <mailto:rafael.lopez@monash.edu>
Here is a bug report concerning (probably) this exact issue: https://tracker.ceph.com/issues/47866 I left a comment describing the situation and my (limited) experiences with it. On 11/11/2020 10:04, Janek Bevendorff wrote:
Yeah, that seems to be it. There are 239 objects prefixed .8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh in my dump. However, there are none of the multiparts from the other file to be found and the head object is 0 bytes.
I checked another multipart object with an end pointer of 11. Surprisingly, it had way more than 11 parts (39 to be precise) named .1, .1_1 .1_2, .1_3, etc. Not sure how Ceph identifies those, but I could find them in the dump at least.
I have no idea why the objects disappeared. I ran a Spark job over all buckets, read 1 byte of every object and recorded errors. Of the 78 buckets, two are missing objects. One bucket is missing one object, the other 15. So, luckily, the incidence is still quite low, but the problem seems to be expanding slowly.
On 10/11/2020 23:46, Rafael Lopez wrote:
Hi Janek,
What you said sounds right - an S3 single part obj won't have an S3 multipart string as part of the prefix. S3 multipart string looks like "2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme".
From memory, single part S3 objects that don't fit in a single rados object are assigned a random prefix that has nothing to do with the object name, and the rados tail/data objects (not the head object) have that prefix. As per your working example, the prefix for that would be '.8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh'. So there would be (239) "shadow" objects with names containing that prefix, and if you add up the sizes it should be the size of your S3 object.
You should look at working and non working examples of both single and multipart S3 objects, as they are probably all a bit different when you look in rados.
I agree it is a serious issue, because once objects are no longer in rados, they cannot be recovered. If it was a case that there was a link broken or rados objects renamed, then we could work to recover...but as far as I can tell, it looks like stuff is just vanishing from rados. The only explanation I can think of is some (rgw or rados) background process is incorrectly doing something with these objects (eg. renaming/deleting). I had thought perhaps it was a bug with the rgw garbage collector..but that is pure speculation.
Once you can articulate the problem, I'd recommend logging a bug tracker upstream.
On Wed, 11 Nov 2020 at 06:33, Janek Bevendorff <janek.bevendorff@uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>> wrote:
Here's something else I noticed: when I stat objects that work via radosgw-admin, the stat info contains a "begin_iter" JSON object with RADOS key info like this
"key": { "name": "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", "instance": "", "ns": "" }
and then "end_iter" with key info like this:
"key": { "name": ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", "instance": "", "ns": "shadow" }
However, when I check the broken 0-byte object, the "begin_iter" and "end_iter" keys look like this:
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", "instance": "", "ns": "multipart" }
[...]
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", "instance": "", "ns": "multipart" }
So, it's the full name plus a suffix and the namespace is multipart, not shadow (or empty). This in itself may just be an artefact of whether the object was uploaded in one go or as a multipart object, but the second difference is that I cannot find any of the multipart objects in my pool's object name dump. I can, however, find the shadow RADOS object of the intact S3 object.
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre
T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> E: rafael.lopez@monash.edu <mailto:rafael.lopez@monash.edu>
This same error caused us to wipe a full cluster of 300TB... will be related to some rados index/database bug not to s3. As Janek exposed is a mayor issue, because the error silent happend and you can only detect it with S3, when you're going to delete/purge a S3 bucket. Dropping NoSuchKey. Error is not related to S3 logic .. Hope this time dev's can take enought time to find and resolve the issue. Error happens with low ec profiles, even with replica x3 in some cases. Regards -----Mensaje original----- De: Janek Bevendorff <janek.bevendorff@uni-weimar.de> Enviado el: jueves, 12 de noviembre de 2020 14:06 Para: Rafael Lopez <rafael.lopez@monash.edu> CC: Robin H. Johnson <robbat2@gentoo.org>; ceph-users <ceph-users@ceph.io> Asunto: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk Here is a bug report concerning (probably) this exact issue: https://tracker.ceph.com/issues/47866 I left a comment describing the situation and my (limited) experiences with it. On 11/11/2020 10:04, Janek Bevendorff wrote:
Yeah, that seems to be it. There are 239 objects prefixed .8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh in my dump. However, there are none of the multiparts from the other file to be found and the head object is 0 bytes.
I checked another multipart object with an end pointer of 11. Surprisingly, it had way more than 11 parts (39 to be precise) named .1, .1_1 .1_2, .1_3, etc. Not sure how Ceph identifies those, but I could find them in the dump at least.
I have no idea why the objects disappeared. I ran a Spark job over all buckets, read 1 byte of every object and recorded errors. Of the 78 buckets, two are missing objects. One bucket is missing one object, the other 15. So, luckily, the incidence is still quite low, but the problem seems to be expanding slowly.
On 10/11/2020 23:46, Rafael Lopez wrote:
Hi Janek,
What you said sounds right - an S3 single part obj won't have an S3 multipart string as part of the prefix. S3 multipart string looks like "2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme".
From memory, single part S3 objects that don't fit in a single rados object are assigned a random prefix that has nothing to do with the object name, and the rados tail/data objects (not the head object) have that prefix. As per your working example, the prefix for that would be '.8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh'. So there would be (239) "shadow" objects with names containing that prefix, and if you add up the sizes it should be the size of your S3 object.
You should look at working and non working examples of both single and multipart S3 objects, as they are probably all a bit different when you look in rados.
I agree it is a serious issue, because once objects are no longer in rados, they cannot be recovered. If it was a case that there was a link broken or rados objects renamed, then we could work to recover...but as far as I can tell, it looks like stuff is just vanishing from rados. The only explanation I can think of is some (rgw or rados) background process is incorrectly doing something with these objects (eg. renaming/deleting). I had thought perhaps it was a bug with the rgw garbage collector..but that is pure speculation.
Once you can articulate the problem, I'd recommend logging a bug tracker upstream.
On Wed, 11 Nov 2020 at 06:33, Janek Bevendorff <janek.bevendorff@uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>> wrote:
Here's something else I noticed: when I stat objects that work via radosgw-admin, the stat info contains a "begin_iter" JSON object with RADOS key info like this
"key": { "name": "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", "instance": "", "ns": "" }
and then "end_iter" with key info like this:
"key": { "name": ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", "instance": "", "ns": "shadow" }
However, when I check the broken 0-byte object, the "begin_iter" and "end_iter" keys look like this:
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", "instance": "", "ns": "multipart" }
[...]
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", "instance": "", "ns": "multipart" }
So, it's the full name plus a suffix and the namespace is multipart, not shadow (or empty). This in itself may just be an artefact of whether the object was uploaded in one go or as a multipart object, but the second difference is that I cannot find any of the multipart objects in my pool's object name dump. I can, however, find the shadow RADOS object of the intact S3 object.
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre
T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> E: rafael.lopez@monash.edu <mailto:rafael.lopez@monash.edu>
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
which Ceph versions are affected by this RGW bug/issues? Luminous, Mimic, Octupos, or the latest? any idea? samuel huxiaoyu@horebdata.cn From: EDH - Manuel Rios Date: 2020-11-12 14:27 To: Janek Bevendorff; Rafael Lopez CC: Robin H. Johnson; ceph-users Subject: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk This same error caused us to wipe a full cluster of 300TB... will be related to some rados index/database bug not to s3. As Janek exposed is a mayor issue, because the error silent happend and you can only detect it with S3, when you're going to delete/purge a S3 bucket. Dropping NoSuchKey. Error is not related to S3 logic .. Hope this time dev's can take enought time to find and resolve the issue. Error happens with low ec profiles, even with replica x3 in some cases. Regards -----Mensaje original----- De: Janek Bevendorff <janek.bevendorff@uni-weimar.de> Enviado el: jueves, 12 de noviembre de 2020 14:06 Para: Rafael Lopez <rafael.lopez@monash.edu> CC: Robin H. Johnson <robbat2@gentoo.org>; ceph-users <ceph-users@ceph.io> Asunto: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk Here is a bug report concerning (probably) this exact issue: https://tracker.ceph.com/issues/47866 I left a comment describing the situation and my (limited) experiences with it. On 11/11/2020 10:04, Janek Bevendorff wrote:
Yeah, that seems to be it. There are 239 objects prefixed .8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh in my dump. However, there are none of the multiparts from the other file to be found and the head object is 0 bytes.
I checked another multipart object with an end pointer of 11. Surprisingly, it had way more than 11 parts (39 to be precise) named .1, .1_1 .1_2, .1_3, etc. Not sure how Ceph identifies those, but I could find them in the dump at least.
I have no idea why the objects disappeared. I ran a Spark job over all buckets, read 1 byte of every object and recorded errors. Of the 78 buckets, two are missing objects. One bucket is missing one object, the other 15. So, luckily, the incidence is still quite low, but the problem seems to be expanding slowly.
On 10/11/2020 23:46, Rafael Lopez wrote:
Hi Janek,
What you said sounds right - an S3 single part obj won't have an S3 multipart string as part of the prefix. S3 multipart string looks like "2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme".
From memory, single part S3 objects that don't fit in a single rados object are assigned a random prefix that has nothing to do with the object name, and the rados tail/data objects (not the head object) have that prefix. As per your working example, the prefix for that would be '.8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh'. So there would be (239) "shadow" objects with names containing that prefix, and if you add up the sizes it should be the size of your S3 object.
You should look at working and non working examples of both single and multipart S3 objects, as they are probably all a bit different when you look in rados.
I agree it is a serious issue, because once objects are no longer in rados, they cannot be recovered. If it was a case that there was a link broken or rados objects renamed, then we could work to recover...but as far as I can tell, it looks like stuff is just vanishing from rados. The only explanation I can think of is some (rgw or rados) background process is incorrectly doing something with these objects (eg. renaming/deleting). I had thought perhaps it was a bug with the rgw garbage collector..but that is pure speculation.
Once you can articulate the problem, I'd recommend logging a bug tracker upstream.
On Wed, 11 Nov 2020 at 06:33, Janek Bevendorff <janek.bevendorff@uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>> wrote:
Here's something else I noticed: when I stat objects that work via radosgw-admin, the stat info contains a "begin_iter" JSON object with RADOS key info like this
"key": { "name": "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", "instance": "", "ns": "" }
and then "end_iter" with key info like this:
"key": { "name": ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", "instance": "", "ns": "shadow" }
However, when I check the broken 0-byte object, the "begin_iter" and "end_iter" keys look like this:
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", "instance": "", "ns": "multipart" }
[...]
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", "instance": "", "ns": "multipart" }
So, it's the full name plus a suffix and the namespace is multipart, not shadow (or empty). This in itself may just be an artefact of whether the object was uploaded in one go or as a multipart object, but the second difference is that I cannot find any of the multipart objects in my pool's object name dump. I can, however, find the shadow RADOS object of the intact S3 object.
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre
T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> E: rafael.lopez@monash.edu <mailto:rafael.lopez@monash.edu>
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I have never seen this on Luminous. I recently upgraded to Octopus and the issue started occurring only few weeks later. On 12/11/2020 16:37, huxiaoyu@horebdata.cn wrote:
which Ceph versions are affected by this RGW bug/issues? Luminous, Mimic, Octupos, or the latest?
any idea?
samuel
------------------------------------------------------------------------ huxiaoyu@horebdata.cn
*From:* EDH - Manuel Rios <mailto:mriosfer@easydatahost.com> *Date:* 2020-11-12 14:27 *To:* Janek Bevendorff <mailto:janek.bevendorff@uni-weimar.de>; Rafael Lopez <mailto:rafael.lopez@monash.edu> *CC:* Robin H. Johnson <mailto:robbat2@gentoo.org>; ceph-users <mailto:ceph-users@ceph.io> *Subject:* [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk This same error caused us to wipe a full cluster of 300TB... will be related to some rados index/database bug not to s3. As Janek exposed is a mayor issue, because the error silent happend and you can only detect it with S3, when you're going to delete/purge a S3 bucket. Dropping NoSuchKey. Error is not related to S3 logic .. Hope this time dev's can take enought time to find and resolve the issue. Error happens with low ec profiles, even with replica x3 in some cases. Regards -----Mensaje original----- De: Janek Bevendorff <janek.bevendorff@uni-weimar.de> Enviado el: jueves, 12 de noviembre de 2020 14:06 Para: Rafael Lopez <rafael.lopez@monash.edu> CC: Robin H. Johnson <robbat2@gentoo.org>; ceph-users <ceph-users@ceph.io> Asunto: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk Here is a bug report concerning (probably) this exact issue: https://tracker.ceph.com/issues/47866 I left a comment describing the situation and my (limited) experiences with it. On 11/11/2020 10:04, Janek Bevendorff wrote: > > Yeah, that seems to be it. There are 239 objects prefixed > .8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh in my dump. However, there are none > of the multiparts from the other file to be found and the head object > is 0 bytes. > > I checked another multipart object with an end pointer of 11. > Surprisingly, it had way more than 11 parts (39 to be precise) named > .1, .1_1 .1_2, .1_3, etc. Not sure how Ceph identifies those, but I > could find them in the dump at least. > > I have no idea why the objects disappeared. I ran a Spark job over all > buckets, read 1 byte of every object and recorded errors. Of the 78 > buckets, two are missing objects. One bucket is missing one object, > the other 15. So, luckily, the incidence is still quite low, but the > problem seems to be expanding slowly. > > > On 10/11/2020 23:46, Rafael Lopez wrote: >> Hi Janek, >> >> What you said sounds right - an S3 single part obj won't have an S3 >> multipart string as part of the prefix. S3 multipart string looks >> like "2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme". >> >> From memory, single part S3 objects that don't fit in a single rados >> object are assigned a random prefix that has nothing to do with >> the object name, and the rados tail/data objects (not the head >> object) have that prefix. >> As per your working example, the prefix for that would be >> '.8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh'. So there would be (239) "shadow" >> objects with names containing that prefix, and if you add up the >> sizes it should be the size of your S3 object. >> >> You should look at working and non working examples of both single >> and multipart S3 objects, as they are probably all a bit different >> when you look in rados. >> >> I agree it is a serious issue, because once objects are no longer in >> rados, they cannot be recovered. If it was a case that there was a >> link broken or rados objects renamed, then we could work to >> recover...but as far as I can tell, it looks like stuff is just >> vanishing from rados. The only explanation I can think of is some >> (rgw or rados) background process is incorrectly doing something with >> these objects (eg. renaming/deleting). I had thought perhaps it was a >> bug with the rgw garbage collector..but that is pure speculation. >> >> Once you can articulate the problem, I'd recommend logging a bug >> tracker upstream. >> >> >> On Wed, 11 Nov 2020 at 06:33, Janek Bevendorff >> <janek.bevendorff@uni-weimar.de >> <mailto:janek.bevendorff@uni-weimar.de>> wrote: >> >> Here's something else I noticed: when I stat objects that work >> via radosgw-admin, the stat info contains a "begin_iter" JSON >> object with RADOS key info like this >> >> >> "key": { >> "name": >> "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", >> "instance": "", >> "ns": "" >> } >> >> >> and then "end_iter" with key info like this: >> >> >> "key": { >> "name": >> ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", >> "instance": "", >> "ns": "shadow" >> } >> >> However, when I check the broken 0-byte object, the "begin_iter" >> and "end_iter" keys look like this: >> >> >> "key": { >> "name": >> "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", >> "instance": "", >> "ns": "multipart" >> } >> >> [...] >> >> >> "key": { >> "name": >> "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", >> "instance": "", >> "ns": "multipart" >> } >> >> So, it's the full name plus a suffix and the namespace is >> multipart, not shadow (or empty). This in itself may just be an >> artefact of whether the object was uploaded in one go or as a >> multipart object, but the second difference is that I cannot find >> any of the multipart objects in my pool's object name dump. I >> can, however, find the shadow RADOS object of the intact S3 object. >> >> >> >> >> -- >> *Rafael Lopez* >> Devops Systems Engineer >> Monash University eResearch Centre >> >> T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> >> E: rafael.lopez@monash.edu <mailto:rafael.lopez@monash.edu> >> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Looks like this is a very dangerous bug for data safety. Hope the bug would be quickly identified and fixed. best regards, Samuel huxiaoyu@horebdata.cn From: Janek Bevendorff Date: 2020-11-12 18:17 To: huxiaoyu@horebdata.cn; EDH - Manuel Rios; Rafael Lopez CC: Robin H. Johnson; ceph-users Subject: Re: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk I have never seen this on Luminous. I recently upgraded to Octopus and the issue started occurring only few weeks later. On 12/11/2020 16:37, huxiaoyu@horebdata.cn wrote: which Ceph versions are affected by this RGW bug/issues? Luminous, Mimic, Octupos, or the latest? any idea? samuel huxiaoyu@horebdata.cn From: EDH - Manuel Rios Date: 2020-11-12 14:27 To: Janek Bevendorff; Rafael Lopez CC: Robin H. Johnson; ceph-users Subject: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk This same error caused us to wipe a full cluster of 300TB... will be related to some rados index/database bug not to s3. As Janek exposed is a mayor issue, because the error silent happend and you can only detect it with S3, when you're going to delete/purge a S3 bucket. Dropping NoSuchKey. Error is not related to S3 logic .. Hope this time dev's can take enought time to find and resolve the issue. Error happens with low ec profiles, even with replica x3 in some cases. Regards -----Mensaje original----- De: Janek Bevendorff <janek.bevendorff@uni-weimar.de> Enviado el: jueves, 12 de noviembre de 2020 14:06 Para: Rafael Lopez <rafael.lopez@monash.edu> CC: Robin H. Johnson <robbat2@gentoo.org>; ceph-users <ceph-users@ceph.io> Asunto: [ceph-users] Re: NoSuchKey on key that is visible in s3 list/radosgw bk Here is a bug report concerning (probably) this exact issue: https://tracker.ceph.com/issues/47866 I left a comment describing the situation and my (limited) experiences with it. On 11/11/2020 10:04, Janek Bevendorff wrote:
Yeah, that seems to be it. There are 239 objects prefixed .8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh in my dump. However, there are none of the multiparts from the other file to be found and the head object is 0 bytes.
I checked another multipart object with an end pointer of 11. Surprisingly, it had way more than 11 parts (39 to be precise) named .1, .1_1 .1_2, .1_3, etc. Not sure how Ceph identifies those, but I could find them in the dump at least.
I have no idea why the objects disappeared. I ran a Spark job over all buckets, read 1 byte of every object and recorded errors. Of the 78 buckets, two are missing objects. One bucket is missing one object, the other 15. So, luckily, the incidence is still quite low, but the problem seems to be expanding slowly.
On 10/11/2020 23:46, Rafael Lopez wrote:
Hi Janek,
What you said sounds right - an S3 single part obj won't have an S3 multipart string as part of the prefix. S3 multipart string looks like "2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme".
From memory, single part S3 objects that don't fit in a single rados object are assigned a random prefix that has nothing to do with the object name, and the rados tail/data objects (not the head object) have that prefix. As per your working example, the prefix for that would be '.8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh'. So there would be (239) "shadow" objects with names containing that prefix, and if you add up the sizes it should be the size of your S3 object.
You should look at working and non working examples of both single and multipart S3 objects, as they are probably all a bit different when you look in rados.
I agree it is a serious issue, because once objects are no longer in rados, they cannot be recovered. If it was a case that there was a link broken or rados objects renamed, then we could work to recover...but as far as I can tell, it looks like stuff is just vanishing from rados. The only explanation I can think of is some (rgw or rados) background process is incorrectly doing something with these objects (eg. renaming/deleting). I had thought perhaps it was a bug with the rgw garbage collector..but that is pure speculation.
Once you can articulate the problem, I'd recommend logging a bug tracker upstream.
On Wed, 11 Nov 2020 at 06:33, Janek Bevendorff <janek.bevendorff@uni-weimar.de <mailto:janek.bevendorff@uni-weimar.de>> wrote:
Here's something else I noticed: when I stat objects that work via radosgw-admin, the stat info contains a "begin_iter" JSON object with RADOS key info like this
"key": { "name": "29/items/WIDE-20110924034843-crawl420/WIDE-20110924065228-02544.warc.gz", "instance": "", "ns": "" }
and then "end_iter" with key info like this:
"key": { "name": ".8naRUHSG2zfgjqmwLnTPvvY1m6DZsgh_239", "instance": "", "ns": "shadow" }
However, when I check the broken 0-byte object, the "begin_iter" and "end_iter" keys look like this:
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.1", "instance": "", "ns": "multipart" }
[...]
"key": { "name": "29/items/WIDE-20110903143858-crawl428/WIDE-20110903143858-01166.warc.gz.2~m5Y42lPMIeis5qgJAZJfuNnzOKd7lme.19", "instance": "", "ns": "multipart" }
So, it's the full name plus a suffix and the namespace is multipart, not shadow (or empty). This in itself may just be an artefact of whether the object was uploaded in one go or as a multipart object, but the second difference is that I cannot find any of the multipart objects in my pool's object name dump. I can, however, find the shadow RADOS object of the intact S3 object.
-- *Rafael Lopez* Devops Systems Engineer Monash University eResearch Centre
T: +61 3 9905 9118 <tel:%2B61%203%209905%209118> E: rafael.lopez@monash.edu <mailto:rafael.lopez@monash.edu>
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
participants (4)
-
EDH - Manuel Rios
-
huxiaoyu@horebdata.cn
-
Janek Bevendorff
-
Rafael Lopez