ceph-users
Threads by month
- ----- 2026 -----
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2025 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2024 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2023 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2022 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2021 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2020 -----
- December
- November
- October
- September
- August
- July
- June
- May
- April
- March
- February
- January
- ----- 2019 -----
- December
- November
- October
- September
- August
- July
- June
- 43 participants
- 9162 discussions
Hi
we're still struggling with our getting our ceph to health_ok. We're
having compounded issues interfering with recovery, as I understand it.
To summarize, we have a cluster of 22 osd nodes running ceph 16.2.x.
About a month back we had one of the OSDs break down (just the OS disk,
but we didn't have a cold spare available, it took a week to get it
fixed). Since the failure of the node, ceph has been repairing the
situation of course, but then it became a problem that our OSDs are
really unevenly balanced (lowest below 50%, highest around 85%). So
whenever a disk fails (and there were 2 since then), the load spreads
over the other OSDs and our fullest OSDs go over the 85% threshold,
slowing down recovery, normal use and rebalancing.
We had issues with degraded PGs, but they weren't being repaired
(because we had turned on the scrubbing during recovery, since we got
messages that lots of PGs weren't being scrubbed in time.
Now there's still one remaining PG degraded because one object is
unfound. The whole error state is taking far too long now and as this is
going on, I was wondering how the balancer wasn't doing its job. Turns
out this is dependent on the cluster being OK or at least not having any
degraded things in it. The balancer hasn't done it's job even though our
cluster was OK for a long time before; we added some 8 nodes a few years
ago and still the newer nodes are having the lowest used OSDs.
Our cluster has about 70-71% usage overall, but with the unbalanced
situation we cannot grow any more. The single node issue (though now
resolved) and ongoing disk failures (we are seeing a handful of OSDs
with read-repaired messages), it looks like we can't get back to health
for a while.
I'm trying to mitigate this by reweighting the fullest OSDs, but the
fuller OSDs keep going over the threshold, while the emptiest OSDs have
plenty of space (just 55% full now).
If you read this far ;-) I'm wondering, can I force repair a PG around
all the restrictions so it doesn't block auto rebalancing?
It seems to me, like that would help, but perhaps there are other things
I can do as well?
(Budget wise, adding more OSD nodes is a bit difficult at the moment...)
Thanks for reading!
Cheers
/Simon
1
0
Introduce: Storage stability testing and DATA consistency verifying tools and system
by 张友加 07 Oct '23
by 张友加 07 Oct '23
07 Oct '23
Dear All,
I hope you are all well. I would like to introduce new tools I have developed, named "LBA tools" which including hd_write_verify & hd_write_verify_dump.
github: https://github.com/zhangyoujia/hd_write_verify
pdf: https://github.com/zhangyoujia/hd_write_verify/DISK&MEMORY stability testing and DATA consistency verifying tools and system.pdf
ppt: https://github.com/zhangyoujia/hd_write_verify/存储稳定性测试与数据一致性校验工具和系统.pptx
bin: https://github.com/zhangyoujia/hd_write_verify/bin
iso: https://github.com/zhangyoujia/hd_write_verify/iso
Data is a vital asset for many businesses, making storage stability and data consistency the most fundamental requirements in storage technology scenarios.
The purpose of storage stability testing is to ensure that storage devices or systems can operate normally and remain stable over time, while also handling various abnormal situations such as sudden power outages and network failures. This testing typically includes stress testing, load testing, fault tolerance testing, and other evaluations to assess the performance and reliability of the storage system.
Data consistency checking is designed to ensure that the data stored in the system is accurate and consistent. This means that whenever data changes occur, all replicas should be updated simultaneously to avoid data inconsistency. Data consistency checking typically involves aspects such as data integrity, accuracy, consistency, and reliability.
LBA tools are very useful for testing Storage stability and verifying DATA consistency, there are much better than FIO & vdbench's verifying functions.
I believe that LBA tools will have a positive impact on the community and help users handle storage data more effectively. Your feedback and suggestions are greatly appreciated, and I hope you can try using LBA tools and share your experiences and recommendations.
Best regards
2
1
Hi
we're still in HEALTH_ERR state with our cluster, this is the top of the
output of `ceph health detail`
HEALTH_ERR 1/846829349 objects unfound (0.000%); 248 scrub errors;
Possible data damage: 1 pg recovery_unfound, 2 pgs inconsistent;
Degraded data redundancy: 6/7118781559 objects degraded (0.000%), 1 pg
degraded, 1 pg undersized; 63 pgs not deep-scrubbed in time; 657 pgs not
scrubbed in time
[WRN] OBJECT_UNFOUND: 1/846829349 objects unfound (0.000%)
pg 26.323 has 1 unfound objects
[ERR] OSD_SCRUB_ERRORS: 248 scrub errors
[ERR] PG_DAMAGED: Possible data damage: 1 pg recovery_unfound, 2 pgs
inconsistent
pg 26.323 is active+recovery_unfound+degraded+remapped, acting
[92,109,116,70,158,128,243,189,256], 1 unfound
pg 26.337 is active+clean+inconsistent, acting
[139,137,48,126,165,89,237,199,189]
pg 26.3e2 is active+clean+inconsistent, acting
[12,27,24,234,195,173,98,32,35]
[WRN] PG_DEGRADED: Degraded data redundancy: 6/7118781559 objects
degraded (0.000%), 1 pg degraded, 1 pg undersized
pg 13.3a5 is stuck undersized for 4m, current state
active+undersized+remapped+backfilling, last acting
[2,45,32,62,2147483647,55,116,25,225,202,240]
pg 26.323 is active+recovery_unfound+degraded+remapped, acting
[92,109,116,70,158,128,243,189,256], 1 unfound
For the PG_DAMAGED pgs I try the usual `ceph pg repair 26.323` etc.,
however it fails to get resolved.
The osd.116 is already marked out and is beginning to get empty. I've
tried restarting the osd processes of the first osd listed for each PG,
but that doesn't get it resolved either.
I guess we should have enough redundancy to get the correct data back,
but how can I tell ceph to fix it in order to get back to a healthy state?
Cheers
/Simon
3
4
Hello All,
Greetings. We've a Ceph Cluster with the version
*ceph version 14.2.16-402-g7d47dbaf4d
(7d47dbaf4d0960a2e910628360ae36def84ed913) nautilus (stable)
===================================
Issues: 1 pg in inconsistent state and does not recover.
# ceph -s
cluster:
id: 30d6f7ee-fa02-4ab3-8a09-9321c8002794
health: HEALTH_ERR
2 large omap objects
1 pools have many more objects per pg than average
159224 scrub errors
Possible data damage: 1 pg inconsistent
2 pgs not deep-scrubbed in time
2 pgs not scrubbed in time
# ceph health detail
HEALTH_ERR 2 large omap objects; 1 pools have many more objects per pg than average; 159224 scrub errors; Possible data damage: 1 pg inconsistent; 2 pgs not deep-scrubbed in time; 2 pgs not scrubbed in time
LARGE_OMAP_OBJECTS 2 large omap objects
2 large objects found in pool 'default.rgw.log'
Search the cluster log for 'Large omap object found' for more details.
MANY_OBJECTS_PER_PG 1 pools have many more objects per pg than average
pool iscsi-images objects per pg (541376) is more than 14.9829 times cluster average (36133)
OSD_SCRUB_ERRORS 159224 scrub errors
PG_DAMAGED Possible data damage: 1 pg inconsistent
pg 15.f4f is active+clean+inconsistent, acting [238,106,402,266,374,498,590,627,684,73,66]
PG_NOT_DEEP_SCRUBBED 2 pgs not deep-scrubbed in time
pg 1.5c not deep-scrubbed since 2021-04-05 23:20:13.714446
pg 1.55 not deep-scrubbed since 2021-04-11 07:12:37.185074
PG_NOT_SCRUBBED 2 pgs not scrubbed in time
pg 1.5c not scrubbed since 2023-07-10 21:15:50.352848
pg 1.55 not scrubbed since 2023-06-24 10:02:10.038311
======================================
We have implemented below command to resolve it
1. We have ran pg repair command "ceph pg repair 15.f4f
2. We have restarted associated OSDs that is mapped to pg 15.f4f
3. We tuned osd_max_scrubs value and set it to 9.
4. We have done scrub and deep scrub by ceph pg scrub 15.4f4 & ceph pg deep-scrub 15.f4f
5. We also tried to ceph-objectstore-tool command to fix it
==============================================
We have checked the logs of the primary OSD of the respective inconsistent PG and found the below errors.
[ERR] : 15.f4fs0 shard 402(2) 15:f2f3fff4:::94a51ddb-a94f-47bc-9068-509e8c09af9a.7862003.20_c%2f4%2fd61%2f885%2f49627697%2f192_1.ts:head : missing
/var/log/ceph/ceph-osd.238.log:339:2023-10-06 00:37:06.410 7f65024cb700 -1 log_channel(cluster) log [ERR] : 15.f4fs0 shard 266(3) 15:f2f00002:::94a51ddb-a94f-47bc-9068-509e8c09af9a.11432468.3_TN8QHE_04.20.2020_08.41%2fCV_MAGNETIC%2fV_274396%2fCHUNK_2440801%2fSFILE_CONTAINER_031.FOLDER%2f3:head : missing
/var/log/ceph/ceph-osd.238.log:340:2023-10-06 00:37:06.410 7f65024cb700 -1 log_channel(cluster) log [ERR] : 15.f4fs0 shard 402(2) 15:f2f00002:::94a51ddb-a94f-47bc-9068-509e8c09af9a.11432468.3_TN8QHE_04.20.2020_08.41%2fCV_MAGNETIC%2fV_274396%2fCHUNK_2440801%2fSFILE_CONTAINER_031.FOLDER%2f3:head : missing
/var/log/ceph/ceph-osd.238.log:341:2023-10-06 00:37:06.410 7f65024cb700 -1 log_channel(cluster) log [ERR] : 15.f4fs0 shard 590(6) 15:f2f00002:::94a51ddb-a94f-47bc-9068-509e8c09af9a.11432468.3_TN8QHE_04.20.2020_08.41%2fCV_MAGNETIC%2fV_274396%2fCHUNK_2440801%2fSFILE_CONTAINER_031.FOLDER%2f3:head : missing
===============================
and also we noticed that the no. of scrub errors in ceph health status are matching with the ERR log entries in the primary OSD logs of the inconsistent PG as below
grep -Hn 'ERR' /var/log/ceph/ceph-osd.238.log|wc -l
159226
================================
Ceph is cleaning the scrub errors but rate of scrub repair is very slow (avg of 200 scrub errors per day) ,we want to increase the rate of scrub error repair to finish the cleanup of pending 159224 scrub errors.
#ceph pg 15.f4f query
{
"state": "active+clean+inconsistent",
"snap_trimq": "[]",
"snap_trimq_len": 0,
"epoch": 409009,
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting_recovery_backfill": [
"66(10)",
"73(9)",
"106(1)",
"238(0)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"info": {
"pgid": "15.f4fs0",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "382701'4900",
"last_user_version": 592883,
"last_backfill": "MAX",
"last_backfill_bitwise": 0,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "409009'7998",
"reported_seq": "379048",
"reported_epoch": "409009",
"state": "active+clean+inconsistent",
"last_fresh": "2023-10-06 15:49:09.174662",
"last_change": "2023-10-06 00:53:53.705308",
"last_active": "2023-10-06 15:49:09.174662",
"last_peered": "2023-10-06 15:49:09.174662",
"last_clean": "2023-10-06 15:49:09.174662",
"last_became_active": "2023-10-03 19:55:56.742034",
"last_became_peered": "2023-10-03 19:55:56.742034",
"last_unstale": "2023-10-06 15:49:09.174662",
"last_undegraded": "2023-10-06 15:49:09.174662",
"last_fullsized": "2023-10-06 15:49:09.174662",
"mapping_epoch": 407096,
"log_start": "382701'4900",
"ondisk_log_start": "382701'4900",
"created": 19813,
"last_epoch_clean": 407097,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3098,
"ondisk_log_size": 3098,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45852508474,
"num_objects": 40234,
"num_object_clones": 0,
"num_object_copies": 442574,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 0,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40234,
"num_whiteouts": 0,
"num_read": 23940,
"num_read_kb": 2900611,
"num_write": 34950,
"num_write_kb": 2440513,
"num_scrub_errors": 159224,
"num_shallow_scrub_errors": 159224,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86741,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [],
"object_location_counts": [],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
"peer_info": [
{
"peer": "66(10)",
"pgid": "15.f4fs10",
"last_update": "409009'7998",
"last_complete": "407068'7791",
"log_tail": "378965'4700",
"last_user_version": 592676,
"last_backfill": "MAX",
"last_backfill_bitwise": 1,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407068'7791",
"reported_seq": "376441",
"reported_epoch": "407068",
"state": "active+clean+inconsistent",
"last_fresh": "2023-10-03 19:07:04.483241",
"last_change": "2023-10-03 18:02:27.182058",
"last_active": "2023-10-03 19:07:04.483241",
"last_peered": "2023-10-03 19:07:04.483241",
"last_clean": "2023-10-03 19:07:04.483241",
"last_became_active": "2023-10-03 18:02:27.181798",
"last_became_peered": "2023-10-03 18:02:27.181798",
"last_unstale": "2023-10-03 19:07:04.483241",
"last_undegraded": "2023-10-03 19:07:04.483241",
"last_fullsized": "2023-10-03 19:07:04.483241",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407068,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3091,
"ondisk_log_size": 3091,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45930070381,
"num_objects": 40379,
"num_object_clones": 0,
"num_object_copies": 444169,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 0,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40379,
"num_whiteouts": 0,
"num_read": 23277,
"num_read_kb": 2877425,
"num_write": 34413,
"num_write_kb": 2409735,
"num_scrub_errors": 0,
"num_shallow_scrub_errors": 0,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [],
"object_location_counts": [],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "73(9)",
"pgid": "15.f4fs9",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 0,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "106(1)",
"pgid": "15.f4fs1",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 1,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "266(3)",
"pgid": "15.f4fs3",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 1,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "374(4)",
"pgid": "15.f4fs4",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 1,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "402(2)",
"pgid": "15.f4fs2",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 0,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "498(5)",
"pgid": "15.f4fs5",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 1,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "590(6)",
"pgid": "15.f4fs6",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 0,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "627(7)",
"pgid": "15.f4fs7",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 1,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
},
{
"peer": "684(8)",
"pgid": "15.f4fs8",
"last_update": "409009'7998",
"last_complete": "409009'7998",
"log_tail": "378965'4700",
"last_user_version": 592677,
"last_backfill": "MAX",
"last_backfill_bitwise": 0,
"purged_snaps": [],
"history": {
"epoch_created": 19813,
"epoch_pool_created": 16141,
"last_epoch_started": 407097,
"last_interval_started": 407096,
"last_epoch_clean": 407097,
"last_interval_clean": 407096,
"last_epoch_split": 19849,
"last_epoch_marked_full": 0,
"same_up_since": 407096,
"same_interval_since": 407096,
"same_primary_since": 407038,
"last_scrub": "408987'7946",
"last_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_deep_scrub": "408987'7946",
"last_deep_scrub_stamp": "2023-10-06 00:53:53.705241",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935"
},
"stats": {
"version": "407095'7792",
"reported_seq": "376677",
"reported_epoch": "407095",
"state": "active+undersized+degraded+inconsistent",
"last_fresh": "2023-10-03 19:54:53.799024",
"last_change": "2023-10-03 19:53:22.457623",
"last_active": "2023-10-03 19:54:53.799024",
"last_peered": "2023-10-03 19:54:53.799024",
"last_clean": "2023-10-03 19:49:47.048960",
"last_became_active": "2023-10-03 19:53:22.457623",
"last_became_peered": "2023-10-03 19:53:22.457623",
"last_unstale": "2023-10-03 19:54:53.799024",
"last_undegraded": "2023-10-03 19:53:22.379335",
"last_fullsized": "2023-10-03 19:53:22.379134",
"mapping_epoch": 407096,
"log_start": "378965'4700",
"ondisk_log_start": "378965'4700",
"created": 19813,
"last_epoch_clean": 407093,
"parent": "0.0",
"parent_split_bits": 0,
"last_scrub": "407037'7783",
"last_scrub_stamp": "2023-10-03 17:14:13.232978",
"last_deep_scrub": "406839'7692",
"last_deep_scrub_stamp": "2023-09-28 14:48:47.098935",
"last_clean_scrub_stamp": "2023-09-28 14:48:47.098935",
"log_size": 3092,
"ondisk_log_size": 3092,
"stats_invalid": false,
"dirty_stats_invalid": false,
"omap_stats_invalid": false,
"hitset_stats_invalid": false,
"hitset_bytes_stats_invalid": false,
"pin_stats_invalid": false,
"manifest_stats_invalid": false,
"snaptrimq_len": 0,
"stat_sum": {
"num_bytes": 45929876177,
"num_objects": 40378,
"num_object_clones": 0,
"num_object_copies": 444158,
"num_objects_missing_on_primary": 0,
"num_objects_missing": 0,
"num_objects_degraded": 40378,
"num_objects_misplaced": 0,
"num_objects_unfound": 0,
"num_objects_dirty": 40378,
"num_whiteouts": 0,
"num_read": 23280,
"num_read_kb": 2877429,
"num_write": 34414,
"num_write_kb": 2409735,
"num_scrub_errors": 159632,
"num_shallow_scrub_errors": 159632,
"num_deep_scrub_errors": 0,
"num_objects_recovered": 86740,
"num_bytes_recovered": 97127524421,
"num_keys_recovered": 0,
"num_objects_omap": 0,
"num_objects_hit_set_archive": 0,
"num_bytes_hit_set_archive": 0,
"num_flush": 0,
"num_flush_kb": 0,
"num_evict": 0,
"num_evict_kb": 0,
"num_promote": 0,
"num_flush_mode_high": 0,
"num_flush_mode_low": 0,
"num_evict_mode_some": 0,
"num_evict_mode_full": 0,
"num_objects_pinned": 0,
"num_legacy_snapsets": 0,
"num_large_omap_objects": 0,
"num_objects_manifest": 0,
"num_omap_bytes": 0,
"num_omap_keys": 0,
"num_objects_repaired": 42593
},
"up": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"acting": [
238,
106,
402,
266,
374,
498,
590,
627,
684,
73,
66
],
"avail_no_missing": [
"238(0)",
"73(9)",
"106(1)",
"266(3)",
"374(4)",
"402(2)",
"498(5)",
"590(6)",
"627(7)",
"684(8)"
],
"object_location_counts": [
{
"shards": "73(9),106(1),238(0),266(3),374(4),402(2),498(5),590(6),627(7),684(8)",
"objects": 40378
}
],
"blocked_by": [],
"up_primary": 238,
"acting_primary": 238,
"purged_snaps": []
},
"empty": 0,
"dne": 0,
"incomplete": 0,
"last_epoch_started": 407097,
"hit_set_history": {
"current_last_update": "0'0",
"history": []
}
}
],
"recovery_state": [
{
"name": "Started/Primary/Active",
"enter_time": "2023-10-03 19:55:56.164849",
"might_have_unfound": [
{
"osd": "66(10)",
"status": "already probed"
},
{
"osd": "73(9)",
"status": "already probed"
},
{
"osd": "106(1)",
"status": "already probed"
},
{
"osd": "266(3)",
"status": "already probed"
},
{
"osd": "374(4)",
"status": "already probed"
},
{
"osd": "402(2)",
"status": "already probed"
},
{
"osd": "498(5)",
"status": "already probed"
},
{
"osd": "590(6)",
"status": "already probed"
},
{
"osd": "627(7)",
"status": "already probed"
},
{
"osd": "684(8)",
"status": "already probed"
}
],
"recovery_progress": {
"backfill_targets": [],
"waiting_on_backfill": [],
"last_backfill_started": "MIN",
"backfill_info": {
"begin": "MIN",
"end": "MIN",
"objects": []
},
"peer_backfill_info": [],
"backfills_in_flight": [],
"recovering": [],
"pg_backend": {
"recovery_ops": [],
"read_ops": []
}
},
"scrub": {
"scrubber.epoch_start": "407096",
"scrubber.active": false,
"scrubber.state": "INACTIVE",
"scrubber.start": "MIN",
"scrubber.end": "MIN",
"scrubber.max_end": "MIN",
"scrubber.subset_last_update": "0'0",
"scrubber.deep": false,
"scrubber.waiting_on_whom": []
}
},
{
"name": "Started",
"enter_time": "2023-10-03 19:55:55.247621"
}
],
"agent_state": {}
}
node1:~ #
1
0
Hi ceph users,
We have a few clusters with quincy 17.2.6 and we are preparing to migrate from ceph-deploy to cephadm for better management.
We are using Ubuntu20 with latest updates (latest openssh).
While testing the migration to cephadm on a test cluster with octopus (v16 latest) we had no issues replacing ceph generated cert/key with our own CA signed certs (ECDSA).
After upgrading to quincy the test cluster and test again the migration we cannot add hosts due to the errors below, ssh access errors specified a while ago in a tracker.
We use the following type of certs:
Type: ecdsa-sha2-nistp384-cert-v01(a)openssh.com user certificate
The certificate works everytime when using ssh client from shell to connect to all hosts in the cluster.
We do a ceph mgr fail every time we replace cert/key so they are restarted.
----- cephadm logs from mgr ------
Oct 06 09:23:27 ceph-m2 bash[1363]: Log: Opening SSH connection to 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Connected to SSH server at 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Local address: 10.10.12.160, port 51870
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Peer address: 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Beginning auth for user root
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Auth failed for user root
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Connection failure: Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Aborting connection
Oct 06 09:23:27 ceph-m2 bash[1363]: Traceback (most recent call last):
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 111, in redirect_log
Oct 06 09:23:27 ceph-m2 bash[1363]: yield
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 90, in _remote_connection
Oct 06 09:23:27 ceph-m2 bash[1363]: preferred_auth=['publickey'], options=ssh_options)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib/python3.6/site-packages/asyncssh/connection.py", line 6804, in connect
Oct 06 09:23:27 ceph-m2 bash[1363]: 'Opening SSH connection to')
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib/python3.6/site-packages/asyncssh/connection.py", line 303, in _connect
Oct 06 09:23:27 ceph-m2 bash[1363]: await conn.wait_established()
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib/python3.6/site-packages/asyncssh/connection.py", line 2243, in wait_established
Oct 06 09:23:27 ceph-m2 bash[1363]: await self._waiter
Oct 06 09:23:27 ceph-m2 bash[1363]: asyncssh.misc.PermissionDenied: Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: During handling of the above exception, another exception occurred:
Oct 06 09:23:27 ceph-m2 bash[1363]: Traceback (most recent call last):
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/orchestrator/_interface.py", line 125, in wrapper
Oct 06 09:23:27 ceph-m2 bash[1363]: return OrchResult(f(*args, **kwargs))
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 2810, in apply
Oct 06 09:23:27 ceph-m2 bash[1363]: results.append(self._apply(spec))
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 2558, in _apply
Oct 06 09:23:27 ceph-m2 bash[1363]: return self._add_host(cast(HostSpec, spec))
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 1434, in _add_host
Oct 06 09:23:27 ceph-m2 bash[1363]: ip_addr = self._check_valid_addr(spec.hostname, spec.addr)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 1415, in _check_valid_addr
Oct 06 09:23:27 ceph-m2 bash[1363]: error_ok=True, no_fsid=True))
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 615, in wait_async
Oct 06 09:23:27 ceph-m2 bash[1363]: return self.event_loop.get_result(coro)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 56, in get_result
Oct 06 09:23:27 ceph-m2 bash[1363]: return asyncio.run_coroutine_threadsafe(coro, self._loop).result()
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib64/python3.6/concurrent/futures/_base.py", line 432, in result
Oct 06 09:23:27 ceph-m2 bash[1363]: return self.__get_result()
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib64/python3.6/concurrent/futures/_base.py", line 384, in __get_result
Oct 06 09:23:27 ceph-m2 bash[1363]: raise self._exception
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/serve.py", line 1361, in _run_cephadm
Oct 06 09:23:27 ceph-m2 bash[1363]: await self.mgr.ssh._remote_connection(host, addr)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 96, in _remote_connection
Oct 06 09:23:27 ceph-m2 bash[1363]: raise
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib64/python3.6/contextlib.py", line 99, in __exit__
Oct 06 09:23:27 ceph-m2 bash[1363]: self.gen.throw(type, value, traceback)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 123, in redirect_log
Oct 06 09:23:27 ceph-m2 bash[1363]: raise HostConnectionError(msg, host, addr)
Oct 06 09:23:27 ceph-m2 bash[1363]: cephadm.ssh.HostConnectionError: Failed to connect to ceph-m1 (10.10.10.232). Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: Log: Opening SSH connection to 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Connected to SSH server at 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Local address: 10.10.12.160, port 51870
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Peer address: 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Beginning auth for user root
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Auth failed for user root
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Connection failure: Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Aborting connection
Oct 06 09:23:27 ceph-m2 bash[1363]: debug 2023-10-06T09:23:27.081+0000 7f78d86d8700 -1 log_channel(cephadm) log [ERR] : Failed to connect to ceph-m1 (10.10.10.232). Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: Log: Opening SSH connection to 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Connected to SSH server at 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Local address: 10.10.12.160, port 51870
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Peer address: 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Beginning auth for user root
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Auth failed for user root
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Connection failure: Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Aborting connection
Oct 06 09:23:27 ceph-m2 bash[1363]: Traceback (most recent call last):
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 111, in redirect_log
Oct 06 09:23:27 ceph-m2 bash[1363]: yield
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 90, in _remote_connection
Oct 06 09:23:27 ceph-m2 bash[1363]: preferred_auth=['publickey'], options=ssh_options)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib/python3.6/site-packages/asyncssh/connection.py", line 6804, in connect
Oct 06 09:23:27 ceph-m2 bash[1363]: 'Opening SSH connection to')
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib/python3.6/site-packages/asyncssh/connection.py", line 303, in _connect
Oct 06 09:23:27 ceph-m2 bash[1363]: await conn.wait_established()
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib/python3.6/site-packages/asyncssh/connection.py", line 2243, in wait_established
Oct 06 09:23:27 ceph-m2 bash[1363]: await self._waiter
Oct 06 09:23:27 ceph-m2 bash[1363]: asyncssh.misc.PermissionDenied: Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: During handling of the above exception, another exception occurred:
Oct 06 09:23:27 ceph-m2 bash[1363]: Traceback (most recent call last):
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/orchestrator/_interface.py", line 125, in wrapper
Oct 06 09:23:27 ceph-m2 bash[1363]: return OrchResult(f(*args, **kwargs))
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 2810, in apply
Oct 06 09:23:27 ceph-m2 bash[1363]: results.append(self._apply(spec))
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 2558, in _apply
Oct 06 09:23:27 ceph-m2 bash[1363]: return self._add_host(cast(HostSpec, spec))
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 1434, in _add_host
Oct 06 09:23:27 ceph-m2 bash[1363]: ip_addr = self._check_valid_addr(spec.hostname, spec.addr)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 1415, in _check_valid_addr
Oct 06 09:23:27 ceph-m2 bash[1363]: error_ok=True, no_fsid=True))
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/module.py", line 615, in wait_async
Oct 06 09:23:27 ceph-m2 bash[1363]: return self.event_loop.get_result(coro)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 56, in get_result
Oct 06 09:23:27 ceph-m2 bash[1363]: return asyncio.run_coroutine_threadsafe(coro, self._loop).result()
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib64/python3.6/concurrent/futures/_base.py", line 432, in result
Oct 06 09:23:27 ceph-m2 bash[1363]: return self.__get_result()
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib64/python3.6/concurrent/futures/_base.py", line 384, in __get_result
Oct 06 09:23:27 ceph-m2 bash[1363]: raise self._exception
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/serve.py", line 1361, in _run_cephadm
Oct 06 09:23:27 ceph-m2 bash[1363]: await self.mgr.ssh._remote_connection(host, addr)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 96, in _remote_connection
Oct 06 09:23:27 ceph-m2 bash[1363]: raise
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/lib64/python3.6/contextlib.py", line 99, in __exit__
Oct 06 09:23:27 ceph-m2 bash[1363]: self.gen.throw(type, value, traceback)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/cephadm/ssh.py", line 123, in redirect_log
Oct 06 09:23:27 ceph-m2 bash[1363]: raise HostConnectionError(msg, host, addr)
Oct 06 09:23:27 ceph-m2 bash[1363]: cephadm.ssh.HostConnectionError: Failed to connect to ceph-m1 (10.10.10.232). Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: Log: Opening SSH connection to 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Connected to SSH server at 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Local address: 10.10.12.160, port 51870
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Peer address: 10.10.10.232, port 22
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Beginning auth for user root
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Auth failed for user root
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Connection failure: Permission denied
Oct 06 09:23:27 ceph-m2 bash[1363]: [conn=3] Aborting connection
Oct 06 09:23:27 ceph-m2 bash[1363]: debug 2023-10-06T09:23:27.081+0000 7f78d86d8700 -1 mgr handle_command module 'orchestrator' command handler threw exception: __init__() missing 2 required positional arguments: >
Oct 06 09:23:27 ceph-m2 bash[1363]: debug 2023-10-06T09:23:27.093+0000 7f78d86d8700 -1 mgr.server reply reply (22) Invalid argument Traceback (most recent call last):
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/mgr_module.py", line 1756, in _handle_command
Oct 06 09:23:27 ceph-m2 bash[1363]: return self.handle_command(inbuf, cmd)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/orchestrator/_interface.py", line 171, in handle_command
Oct 06 09:23:27 ceph-m2 bash[1363]: return dispatch[cmd['prefix']].call(self, cmd, inbuf)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/mgr_module.py", line 462, in call
Oct 06 09:23:27 ceph-m2 bash[1363]: return self.func(mgr, **kwargs)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/orchestrator/_interface.py", line 107, in <lambda>
Oct 06 09:23:27 ceph-m2 bash[1363]: wrapper_copy = lambda *l_args, **l_kwargs: wrapper(*l_args, **l_kwargs) # noqa: E731
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/orchestrator/_interface.py", line 96, in wrapper
Oct 06 09:23:27 ceph-m2 bash[1363]: return func(*args, **kwargs)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/orchestrator/module.py", line 356, in _add_host
Oct 06 09:23:27 ceph-m2 bash[1363]: return self._apply_misc([s], False, Format.plain)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/orchestrator/module.py", line 1092, in _apply_misc
Oct 06 09:23:27 ceph-m2 bash[1363]: raise_if_exception(completion)
Oct 06 09:23:27 ceph-m2 bash[1363]: File "/usr/share/ceph/mgr/orchestrator/_interface.py", line 225, in raise_if_exception
Oct 06 09:23:27 ceph-m2 bash[1363]: e = pickle.loads(c.serialized_exception)
Oct 06 09:23:27 ceph-m2 bash[1363]: TypeError: __init__() missing 2 required positional arguments: 'hostname' and 'addr'
----- cephadm logs from mgr ------
----- sshd logs DEBUG3 level ------
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug2: input_userauth_request: try method publickey [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug2: userauth_pubkey: valid user root querying public key ecdsa-sha2-nistp384 AAAAE2VjZHNhLXNoYTItbmlzdHAzO------------ [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: userauth_pubkey: test pkalg ecdsa-sha2-nistp384 pkblob ECDSA SHA256:m6Q0ZQVjjDLWxbmCn0hcGQ2---------- [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_key_allowed entering [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_request_send entering: type 22 [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_key_allowed: waiting for MONITOR_ANS_KEYALLOWED [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_request_receive_expect entering: type 23 [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_request_receive entering [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_request_receive entering
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: monitor_read: checking request 22
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_answer_keyallowed entering
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_answer_keyallowed: key_from_blob: 0x5568f0aa7880
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: temporarily_use_uid: 0/0 (e=0/0)
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: trying public key file /etc/ssh/fake_authorized_keys
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: fd 5 clearing O_NONBLOCK
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: restore_uid: 0/0
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_answer_keyallowed: publickey authentication test: ECDSA key is not allowed
Oct 6 09:33:09 ceph-m1 sshd[57168]: Failed publickey for root from 10.10.12.160 port 40854 ssh2: ECDSA SHA256:m6Q0ZQVjjDLWxbmCn0hcGQ24gbpk-------------
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_request_send entering: type 23
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug2: userauth_pubkey: authenticated 0 pkalg ecdsa-sha2-nistp384 [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: user_specific_delay: user specific delay 0.000ms [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: ensure_minimum_time_since: elapsed 8.263ms, delaying 8.080ms (requested 8.171ms) [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: userauth_finish: failure partial=0 next methods="publickey" [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: send packet: type 51 [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: Connection closed by authenticating user root 10.10.12.160 port 40854 [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: do_cleanup [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: PAM: sshpam_thread_cleanup entering [preauth]
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: monitor_read_log: child log fd closed
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: mm_request_receive entering
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: do_cleanup
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: PAM: cleanup
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug3: PAM: sshpam_thread_cleanup entering
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: Killing privsep child 57169
Oct 6 09:33:09 ceph-m1 sshd[57168]: debug1: audit_event: unhandled event 12
Oct 6 09:33:09 ceph-m1 sshd[757]: debug1: main_sigchld_handler: Child exited
---------------
I get "ECDSA key is not allowed" above.
From sshd logs, it looks like the client is not sending what is required or in the expected format.
Now, what was changed in quincy/mgr on ssh client?
Is anyone else using ECDSA keys and it works with quincy?
I could not find in PRs something specific to this that could block the access, but it might be.
Any suggestion?
Thank you!
Paul
1
0
06 Oct '23
Hello
Short question regarding journal-based rbd mirroring.
▪IO path with journaling w/o cache:
a. Create an event to describe the update
b. Asynchronously append event to journal object
c. Asynchronously update image once event is safe
d. Complete IO to client once update is safe
[cf. https://events.static.linuxfound.org/sites/events/files/slides/Disaster%20R…]
If a client crashes between b. and c., is there a mechanism to replay the IO from the journal on the primary image?
If not, then the primary and secondary images would get out-of-sync (because of the extra write(s) on secondary) and subsequent writes to the primary would corrupt the secondary. Is that correct?
Cheers
Francois Scheurer
--
EveryWare AG
François Scheurer
Senior Systems Engineer
Zurlindenstrasse 52a
CH-8003 Zürich
tel: +41 44 466 60 00
fax: +41 44 466 60 10
mail: francois.scheurer(a)everyware.ch
web: http://www.everyware.ch
1
1
Hi,
I am still evaluating ceph rgw for specific use cases.
My question is about keeping the realm of bucket names under control of
rgw admins.
Normal S3 users have the ability to create new buckets as they see fit.
This opens opportunities for creating excessive amounts of buckets, or
for blocking nice bucket names for other uses, or even using
bucketname-typosquatting as an attack vector.
In AWS, I can create some IAM users and provide per-bucket access to
them via bucket or IAM user policies. These IAM users can't create new
buckets on their own. Giving out only those IAM credentials to users and
applications, I can ensure no bucket namespace pollution occurs.
Ceph rgw does not have IAM users (yet?). What could I use here to not
allow certain S3 users to create buckets on their own?
Regards
Matthias
5
10
Hi,
I've just upgraded to our object storages to the latest pacific version
(16.2.14) and the autscaler is acting weird.
On one cluster it just shows nothing:
~# ceph osd pool autoscale-status
~#
On the other clusters it shows this when it is set to warn:
~# ceph health detail
...
[WRN] POOL_TOO_MANY_PGS: 2 pools have too many placement groups
Pool .rgw.buckets.data has 1024 placement groups, should have 1024
Pool device_health_metrics has 1 placement groups, should have 1
Version 16.2.13 seems to act normal.
Is this a known bug?
--
Die Selbsthilfegruppe "UTF-8-Probleme" trifft sich diesmal abweichend im
groüen Saal.
2
4
I have an 8-node cluster with old hardware. a week ago 4 nodes went down and the CEPH cluster went nuts.
All pgs became unknown and montors took too long to be in sync.
So i reduced the number of mons to one and mgrs to one as well
Now the recovery starts with 100% unknown pgs and then pgs start to move ot inactive . It generally fails to recover in the middle and starts from scratch.
It's hold hardware and OSDs have lots of slow ops and probably number of bad sectors as well
Any suggestions on how to tackle this. It's a nautilus cluster and pretty old (8-year old hardware)
Thanks
3
2
Hi Team,Milind
*Ceph-version:* Quincy, Reef
*OS:* Almalinux 8
*Issue:* snap_schedule works after 1 hour of schedule
*Description:*
We are currently working in a 3-node ceph cluster.
We are currently exploring the scheduled snapshot capability of the
ceph-mgr module.
To enable/configure scheduled snapshots, we followed the following link:
https://docs.ceph.com/en/quincy/cephfs/snap-schedule/
We were able to create snap schedules for the subvolumes as suggested.
But we have observed a two very strange behaviour:
1. The snap_schedules only work when we restart the ceph-mgr service on the
mgr node:
We then restarted the mgr-service on the active mgr node, and after 1 hour
it started getting created. I am attaching the log file for the same after
restart. Thre behaviour looks abnormal.
So, for eg consider the below output:
```
[root@storagenode-1 ~]# ceph fs snap-schedule status
/volumes/subvolgrp/test3
{"fs": "cephfs", "subvol": null, "path": "/volumes/subvolgrp/test3",
"rel_path": "/volumes/subvolgrp/test3", "schedule": "1h", "retention": {},
"start": "2023-10-04T07:20:00", "created": "2023-10-04T07:18:41", "first":
"2023-10-04T08:20:00", "last": "2023-10-04T09:20:00", "last_pruned": null,
"created_count": 2, "pruned_count": 0, "active": true}
[root@storagenode-1 ~]#
```
As we can see in the above o/p, we created the schedule at
2023-10-04T07:18:41. The schedule was suppose to start at
2023-10-04T07:20:00 but it started at 2023-10-04T08:20:00
Any input w.r.t the same will be of great help.
Thanks and Regards
Kushagra Gupta
2
7