Ceph 16.2.x mon compactions, disk writes
Hi, Monitors in our 16.2.14 cluster appear to quite often run "manual compaction" tasks: debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", "num_memtables": 1, "num_entries": 715, "num_deletes": 251, "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) I have removed a lot of interim log messages to save space. During each compaction the monitor process writes approximately 500-600 MB of data to disk over a short period of time. These writes add up to tens of gigabytes per hour and hundreds of gigabytes per day. Monitor rocksdb and compaction options are default: "mon_compact_on_bootstrap": "false", "mon_compact_on_start": "false", "mon_compact_on_trim": "true", "mon_rocksdb_options": "write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true", Is this expected behavior? Is this something I can adjust in order to extend the system storage life? Best regards, Zakhar
Any input from anyone, please? It's another thing that seems to be rather poorly documented: it's unclear what to expect, what 'normal' behavior should be, and what can be done about the huge amount of writes by monitors. /Z On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Hi,
Monitors in our 16.2.14 cluster appear to quite often run "manual compaction" tasks:
debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", "num_memtables": 1, "num_entries": 715, "num_deletes": 251, "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end)
I have removed a lot of interim log messages to save space.
During each compaction the monitor process writes approximately 500-600 MB of data to disk over a short period of time. These writes add up to tens of gigabytes per hour and hundreds of gigabytes per day.
Monitor rocksdb and compaction options are default:
"mon_compact_on_bootstrap": "false", "mon_compact_on_start": "false", "mon_compact_on_trim": "true", "mon_rocksdb_options": "write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
Is this expected behavior? Is this something I can adjust in order to extend the system storage life?
Best regards, Zakhar
Any input from anyone, please? On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Any input from anyone, please?
It's another thing that seems to be rather poorly documented: it's unclear what to expect, what 'normal' behavior should be, and what can be done about the huge amount of writes by monitors.
/Z
On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Hi,
Monitors in our 16.2.14 cluster appear to quite often run "manual compaction" tasks:
debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", "num_memtables": 1, "num_entries": 715, "num_deletes": 251, "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end)
I have removed a lot of interim log messages to save space.
During each compaction the monitor process writes approximately 500-600 MB of data to disk over a short period of time. These writes add up to tens of gigabytes per hour and hundreds of gigabytes per day.
Monitor rocksdb and compaction options are default:
"mon_compact_on_bootstrap": "false", "mon_compact_on_start": "false", "mon_compact_on_trim": "true", "mon_rocksdb_options": "write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
Is this expected behavior? Is this something I can adjust in order to extend the system storage life?
Best regards, Zakhar
Hi, what you report is the expected behaviour, at least I see the same on all clusters. I can't answer why the compaction is required that often, but you can control the log level of the rocksdb output: ceph config set mon debug_rocksdb 1/5 (default is 4/5) This reduces the log entries and you wouldn't see the manual compaction logs anymore. There are a couple more rocksdb options but I probably wouldn't change too much, only if you know what you're doing. Maybe Igor can comment if some other tuning makes sense here. Regards, Eugen Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Any input from anyone, please?
On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Any input from anyone, please?
It's another thing that seems to be rather poorly documented: it's unclear what to expect, what 'normal' behavior should be, and what can be done about the huge amount of writes by monitors.
/Z
On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Hi,
Monitors in our 16.2.14 cluster appear to quite often run "manual compaction" tasks:
debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", "num_memtables": 1, "num_entries": 715, "num_deletes": 251, "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end)
I have removed a lot of interim log messages to save space.
During each compaction the monitor process writes approximately 500-600 MB of data to disk over a short period of time. These writes add up to tens of gigabytes per hour and hundreds of gigabytes per day.
Monitor rocksdb and compaction options are default:
"mon_compact_on_bootstrap": "false", "mon_compact_on_start": "false", "mon_compact_on_trim": "true", "mon_rocksdb_options": "write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
Is this expected behavior? Is this something I can adjust in order to extend the system storage life?
Best regards, Zakhar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Thank you, Eugen. I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible. /Z On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
Hi,
what you report is the expected behaviour, at least I see the same on all clusters. I can't answer why the compaction is required that often, but you can control the log level of the rocksdb output:
ceph config set mon debug_rocksdb 1/5 (default is 4/5)
This reduces the log entries and you wouldn't see the manual compaction logs anymore. There are a couple more rocksdb options but I probably wouldn't change too much, only if you know what you're doing. Maybe Igor can comment if some other tuning makes sense here.
Regards, Eugen
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Any input from anyone, please?
On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Any input from anyone, please?
It's another thing that seems to be rather poorly documented: it's unclear what to expect, what 'normal' behavior should be, and what can be done about the huge amount of writes by monitors.
/Z
On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Hi,
Monitors in our 16.2.14 cluster appear to quite often run "manual compaction" tasks:
debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", "num_memtables": 1, "num_entries": 715, "num_deletes": 251, "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end)
I have removed a lot of interim log messages to save space.
During each compaction the monitor process writes approximately 500-600 MB of data to disk over a short period of time. These writes add up to tens of gigabytes per hour and hundreds of gigabytes per day.
Monitor rocksdb and compaction options are default:
"mon_compact_on_bootstrap": "false", "mon_compact_on_start": "false", "mon_compact_on_trim": "true", "mon_rocksdb_options":
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
Is this expected behavior? Is this something I can adjust in order to extend the system storage life?
Best regards, Zakhar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I need to ask here: where exactly do you observe the hundreds of GB written per day? Are the mon logs huge? Is it the mon store? Is your cluster unhealthy? We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about 800 librbd clients. Per week our mon logs are about 70M, the cluster logs about 120M , the audit logs about 70M and I see between 100-200Kb/s writes to the mon store. That's in the lower-digit GB range per day. Hundreds of GB per day sound completely over the top on a healthy cluster, unless you have MGR modules changing the OSD/cluster map continuously. Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed by now? @Eugen: Was there not an option to limit logging to the MON store? For information to readers, we followed old recommendations from a Dell white paper for building a ceph cluster and have a 1TB Raid10 array on 6x write intensive SSDs for the MON stores. After 5 years we are below 10% wear. Average size of the MON store for a healthy cluster is 500M-1G, but we have seen this ballooning to 100+GB in degraded conditions. Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14 ________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 12:00 PM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes Thank you, Eugen. I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible. /Z On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
Hi,
what you report is the expected behaviour, at least I see the same on all clusters. I can't answer why the compaction is required that often, but you can control the log level of the rocksdb output:
ceph config set mon debug_rocksdb 1/5 (default is 4/5)
This reduces the log entries and you wouldn't see the manual compaction logs anymore. There are a couple more rocksdb options but I probably wouldn't change too much, only if you know what you're doing. Maybe Igor can comment if some other tuning makes sense here.
Regards, Eugen
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Any input from anyone, please?
On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Any input from anyone, please?
It's another thing that seems to be rather poorly documented: it's unclear what to expect, what 'normal' behavior should be, and what can be done about the huge amount of writes by monitors.
/Z
On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Hi,
Monitors in our 16.2.14 cluster appear to quite often run "manual compaction" tasks:
debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", "num_memtables": 1, "num_entries": 715, "num_deletes": 251, "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end)
I have removed a lot of interim log messages to save space.
During each compaction the monitor process writes approximately 500-600 MB of data to disk over a short period of time. These writes add up to tens of gigabytes per hour and hundreds of gigabytes per day.
Monitor rocksdb and compaction options are default:
"mon_compact_on_bootstrap": "false", "mon_compact_on_start": "false", "mon_compact_on_trim": "true", "mon_rocksdb_options":
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
Is this expected behavior? Is this something I can adjust in order to extend the system storage life?
Best regards, Zakhar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Thank you, Frank. The cluster is healthy, operating normally, nothing unusual is going on. We observe lots of writes by mon processes into mon rocksdb stores, specifically: /var/lib/ceph/mon/ceph-cephXX/store.db: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst The site of the files is not huge, but monitors rotate and write out these files often, sometimes several times per minute, resulting in lots of data written to disk. The writes coincide with "manual compaction" events logged by the monitors, for example: debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1676] [default] [JOB 70854] Compacting 1@5 + 9@6 files to L6, score -1.00 debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1697022610487624, "job": 70854, "event": "compaction_started", "compaction_reason": "ManualCompaction", "files_L5": [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536, 3675537, 3675538, 3675539, 3675540, 3675541], "score": -1, "input_data_size": 601117031} debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675544: 2015 keys, 67287115 bytes debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675545: 24343 keys, 67336225 bytes debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675546: 1196 keys, 67225813 bytes debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675547: 1049 keys, 67252678 bytes debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675548: 1081 keys, 67216638 bytes debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675549: 1196 keys, 67245376 bytes debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675550: 1195 keys, 67246813 bytes debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675551: 1205 keys, 67223302 bytes debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675552: 1312 keys, 56416011 bytes debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted 1@5 + 9@6 files to L6 => 594449971 bytes debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760] [default] compacted to: base level 5 level multiplier 10.00 max bytes base 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd, 173.9 wr, level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9), read-write-amplify(3434.6) write-amplify(1707.7) OK, records in: 35108, records dropped: 516 output_compression: NoCompression debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros": 1697022613921002, "job": 70854, "event": "compaction_finished", "compaction_time_micros": 3418822, "compaction_time_cpu_micros": 785454, "output_level": 6, "num_output_files": 9, "total_output_size": 594449971, "num_input_records": 35108, "num_output_records": 34592, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 0, 0, 0, 0, 0, 9]} The log even mentions the huge write multiplication. I wonder whether this is normal and what can be done about it. /Z On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk> wrote:
I need to ask here: where exactly do you observe the hundreds of GB written per day? Are the mon logs huge? Is it the mon store? Is your cluster unhealthy?
We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about 800 librbd clients. Per week our mon logs are about 70M, the cluster logs about 120M , the audit logs about 70M and I see between 100-200Kb/s writes to the mon store. That's in the lower-digit GB range per day. Hundreds of GB per day sound completely over the top on a healthy cluster, unless you have MGR modules changing the OSD/cluster map continuously.
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed by now?
@Eugen: Was there not an option to limit logging to the MON store?
For information to readers, we followed old recommendations from a Dell white paper for building a ceph cluster and have a 1TB Raid10 array on 6x write intensive SSDs for the MON stores. After 5 years we are below 10% wear. Average size of the MON store for a healthy cluster is 500M-1G, but we have seen this ballooning to 100+GB in degraded conditions.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 12:00 PM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Thank you, Eugen.
I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible.
/Z
On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
Hi,
what you report is the expected behaviour, at least I see the same on all clusters. I can't answer why the compaction is required that often, but you can control the log level of the rocksdb output:
ceph config set mon debug_rocksdb 1/5 (default is 4/5)
This reduces the log entries and you wouldn't see the manual compaction logs anymore. There are a couple more rocksdb options but I probably wouldn't change too much, only if you know what you're doing. Maybe Igor can comment if some other tuning makes sense here.
Regards, Eugen
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Any input from anyone, please?
On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Any input from anyone, please?
It's another thing that seems to be rather poorly documented: it's unclear what to expect, what 'normal' behavior should be, and what can be done about the huge amount of writes by monitors.
/Z
On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Hi,
Monitors in our 16.2.14 cluster appear to quite often run "manual compaction" tasks:
debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", "num_memtables": 1, "num_entries": 715, "num_deletes": 251, "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end)
I have removed a lot of interim log messages to save space.
During each compaction the monitor process writes approximately 500-600 MB of data to disk over a short period of time. These writes add up to tens of gigabytes per hour and hundreds of gigabytes per day.
Monitor rocksdb and compaction options are default:
"mon_compact_on_bootstrap": "false", "mon_compact_on_start": "false", "mon_compact_on_trim": "true", "mon_rocksdb_options":
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
Is this expected behavior? Is this something I can adjust in order to extend the system storage life?
Best regards, Zakhar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Do you use many snapshots (rbd or cephfs)? That can cause a heavy monitor usage, we've seen large mon stores on customer clusters with rbd mirroring on snapshot basis. In a healthy cluster they have mon stores of around 2GB in size.
@Eugen: Was there not an option to limit logging to the MON store?
I don't recall at the moment, worth checking tough. Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Thank you, Frank.
The cluster is healthy, operating normally, nothing unusual is going on. We observe lots of writes by mon processes into mon rocksdb stores, specifically:
/var/lib/ceph/mon/ceph-cephXX/store.db: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
The site of the files is not huge, but monitors rotate and write out these files often, sometimes several times per minute, resulting in lots of data written to disk. The writes coincide with "manual compaction" events logged by the monitors, for example:
debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1676] [default] [JOB 70854] Compacting 1@5 + 9@6 files to L6, score -1.00 debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1697022610487624, "job": 70854, "event": "compaction_started", "compaction_reason": "ManualCompaction", "files_L5": [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536, 3675537, 3675538, 3675539, 3675540, 3675541], "score": -1, "input_data_size": 601117031} debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675544: 2015 keys, 67287115 bytes debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675545: 24343 keys, 67336225 bytes debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675546: 1196 keys, 67225813 bytes debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675547: 1049 keys, 67252678 bytes debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675548: 1081 keys, 67216638 bytes debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675549: 1196 keys, 67245376 bytes debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675550: 1195 keys, 67246813 bytes debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675551: 1205 keys, 67223302 bytes debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675552: 1312 keys, 56416011 bytes debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted 1@5 + 9@6 files to L6 => 594449971 bytes debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760] [default] compacted to: base level 5 level multiplier 10.00 max bytes base 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd, 173.9 wr, level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9), read-write-amplify(3434.6) write-amplify(1707.7) OK, records in: 35108, records dropped: 516 output_compression: NoCompression debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros": 1697022613921002, "job": 70854, "event": "compaction_finished", "compaction_time_micros": 3418822, "compaction_time_cpu_micros": 785454, "output_level": 6, "num_output_files": 9, "total_output_size": 594449971, "num_input_records": 35108, "num_output_records": 34592, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 0, 0, 0, 0, 0, 9]}
The log even mentions the huge write multiplication. I wonder whether this is normal and what can be done about it.
/Z
On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk> wrote:
I need to ask here: where exactly do you observe the hundreds of GB written per day? Are the mon logs huge? Is it the mon store? Is your cluster unhealthy?
We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about 800 librbd clients. Per week our mon logs are about 70M, the cluster logs about 120M , the audit logs about 70M and I see between 100-200Kb/s writes to the mon store. That's in the lower-digit GB range per day. Hundreds of GB per day sound completely over the top on a healthy cluster, unless you have MGR modules changing the OSD/cluster map continuously.
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed by now?
@Eugen: Was there not an option to limit logging to the MON store?
For information to readers, we followed old recommendations from a Dell white paper for building a ceph cluster and have a 1TB Raid10 array on 6x write intensive SSDs for the MON stores. After 5 years we are below 10% wear. Average size of the MON store for a healthy cluster is 500M-1G, but we have seen this ballooning to 100+GB in degraded conditions.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 12:00 PM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Thank you, Eugen.
I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible.
/Z
On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
Hi,
what you report is the expected behaviour, at least I see the same on all clusters. I can't answer why the compaction is required that often, but you can control the log level of the rocksdb output:
ceph config set mon debug_rocksdb 1/5 (default is 4/5)
This reduces the log entries and you wouldn't see the manual compaction logs anymore. There are a couple more rocksdb options but I probably wouldn't change too much, only if you know what you're doing. Maybe Igor can comment if some other tuning makes sense here.
Regards, Eugen
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Any input from anyone, please?
On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Any input from anyone, please?
It's another thing that seems to be rather poorly documented: it's unclear what to expect, what 'normal' behavior should be, and what can be done about the huge amount of writes by monitors.
/Z
On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Hi,
Monitors in our 16.2.14 cluster appear to quite often run "manual compaction" tasks:
debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", "num_memtables": 1, "num_entries": 715, "num_deletes": 251, "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; will stop at (end) debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": "Manual Compaction"} debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; will stop at (end) debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction starting debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; will stop at (end)
I have removed a lot of interim log messages to save space.
During each compaction the monitor process writes approximately 500-600 MB of data to disk over a short period of time. These writes add up to tens of gigabytes per hour and hundreds of gigabytes per day.
Monitor rocksdb and compaction options are default:
"mon_compact_on_bootstrap": "false", "mon_compact_on_start": "false", "mon_compact_on_trim": "true", "mon_rocksdb_options":
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
Is this expected behavior? Is this something I can adjust in order to extend the system storage life?
Best regards, Zakhar
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
We don't use CephFS at all and don't have RBD snapshots apart from some cloning for Openstack images. The size of mon stores isn't an issue, it's < 600 MB. But it gets overwritten often causing lots of disk writes, and that is an issue for us. /Z On Wed, 11 Oct 2023 at 14:37, Eugen Block <eblock@nde.ag> wrote:
Do you use many snapshots (rbd or cephfs)? That can cause a heavy monitor usage, we've seen large mon stores on customer clusters with rbd mirroring on snapshot basis. In a healthy cluster they have mon stores of around 2GB in size.
@Eugen: Was there not an option to limit logging to the MON store?
I don't recall at the moment, worth checking tough.
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
9@6 files to L6 => 594449971 bytes debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760] [default] compacted to: base level 5 level multiplier 10.00 max bytes base 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd, 173.9 wr, level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9), read-write-amplify(3434.6) write-amplify(1707.7) OK, records in: 35108, records dropped: 516 output_compression: NoCompression debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros": 1697022613921002, "job": 70854, "event": "compaction_finished", "compaction_time_micros": 3418822, "compaction_time_cpu_micros": 785454, "output_level": 6, "num_output_files": 9, "total_output_size": 594449971, "num_input_records": 35108, "num_output_records": 34592, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 0, 0, 0, 0, 0, 9]}
The log even mentions the huge write multiplication. I wonder whether this is normal and what can be done about it.
/Z
On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk> wrote:
I need to ask here: where exactly do you observe the hundreds of GB written per day? Are the mon logs huge? Is it the mon store? Is your cluster unhealthy?
We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about 800 librbd clients. Per week our mon logs are about 70M, the cluster logs about 120M , the audit logs about 70M and I see between 100-200Kb/s writes to the mon store. That's in the lower-digit GB range per day. Hundreds of GB per day sound completely over the top on a healthy cluster, unless you have MGR modules changing the OSD/cluster map continuously.
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed by now?
@Eugen: Was there not an option to limit logging to the MON store?
For information to readers, we followed old recommendations from a Dell white paper for building a ceph cluster and have a 1TB Raid10 array on 6x write intensive SSDs for the MON stores. After 5 years we are below 10% wear. Average size of the MON store for a healthy cluster is 500M-1G, but we have seen this ballooning to 100+GB in degraded conditions.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 12:00 PM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Thank you, Eugen.
I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible.
/Z
On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
Hi,
what you report is the expected behaviour, at least I see the same on all clusters. I can't answer why the compaction is required that often, but you can control the log level of the rocksdb output:
ceph config set mon debug_rocksdb 1/5 (default is 4/5)
This reduces the log entries and you wouldn't see the manual compaction logs anymore. There are a couple more rocksdb options but I probably wouldn't change too much, only if you know what you're doing. Maybe Igor can comment if some other tuning makes sense here.
Regards, Eugen
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Any input from anyone, please?
On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Any input from anyone, please?
It's another thing that seems to be rather poorly documented: it's unclear what to expect, what 'normal' behavior should be, and what can be done about the huge amount of writes by monitors.
/Z
On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
> Hi, > > Monitors in our 16.2.14 cluster appear to quite often run "manual > compaction" tasks: > > debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 > {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", > "num_memtables": 1, "num_entries": 715, "num_deletes": 251, > "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": > "Manual Compaction"} > debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log > Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] > [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; > will stop at (end) > debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log > Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] > [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; > will stop at (end) > debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 > {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", > "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, > "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": > "Manual Compaction"} > debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log > Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] > [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; > will stop at (end) > debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log > Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] > [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; > will stop at (end) > debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 > {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", > "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, > "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": > "Manual Compaction"} > debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log > Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] > [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; > will stop at (end) > debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > starting > debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log > Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] > [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; > will stop at (end) > > I have removed a lot of interim log messages to save space. > > During each compaction the monitor process writes approximately 500-600 > MB of data to disk over a short period of time. These writes add up to tens > of gigabytes per hour and hundreds of gigabytes per day. > > Monitor rocksdb and compaction options are default: > > "mon_compact_on_bootstrap": "false", > "mon_compact_on_start": "false", > "mon_compact_on_trim": "true", > "mon_rocksdb_options": >
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
> > Is this expected behavior? Is this something I can adjust in order to > extend the system storage life? > > Best regards, > Zakhar >
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
9@6 files to L6, score -1.00 debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1697022610487624, "job": 70854, "event": "compaction_started", "compaction_reason": "ManualCompaction", "files_L5": [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536, 3675537, 3675538, 3675539, 3675540, 3675541], "score": -1, "input_data_size": 601117031} debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675544: 2015 keys, 67287115 bytes debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675545: 24343 keys, 67336225 bytes debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675546: 1196 keys, 67225813 bytes debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675547: 1049 keys, 67252678 bytes debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675548: 1081 keys, 67216638 bytes debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675549: 1196 keys, 67245376 bytes debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675550: 1195 keys, 67246813 bytes debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675551: 1205 keys, 67223302 bytes debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675552: 1312 keys, 56416011 bytes debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted 1@5
Thank you, Frank.
The cluster is healthy, operating normally, nothing unusual is going on. We observe lots of writes by mon processes into mon rocksdb stores, specifically:
/var/lib/ceph/mon/ceph-cephXX/store.db: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
The site of the files is not huge, but monitors rotate and write out these files often, sometimes several times per minute, resulting in lots of data written to disk. The writes coincide with "manual compaction" events logged by the monitors, for example:
debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1676] [default] [JOB 70854] Compacting 1@5
Can you add some more details as requested by Frank? Which mgr modules are enabled? What's the current 'ceph -s' output?
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
We don't use CephFS at all and don't have RBD snapshots apart from some cloning for Openstack images.
The size of mon stores isn't an issue, it's < 600 MB. But it gets overwritten often causing lots of disk writes, and that is an issue for us.
/Z
On Wed, 11 Oct 2023 at 14:37, Eugen Block <eblock@nde.ag> wrote:
Do you use many snapshots (rbd or cephfs)? That can cause a heavy monitor usage, we've seen large mon stores on customer clusters with rbd mirroring on snapshot basis. In a healthy cluster they have mon stores of around 2GB in size.
@Eugen: Was there not an option to limit logging to the MON store?
I don't recall at the moment, worth checking tough.
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
9@6 files to L6 => 594449971 bytes debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760] [default] compacted to: base level 5 level multiplier 10.00 max bytes base 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd, 173.9 wr, level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9), read-write-amplify(3434.6) write-amplify(1707.7) OK, records in: 35108, records dropped: 516 output_compression: NoCompression debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros": 1697022613921002, "job": 70854, "event": "compaction_finished", "compaction_time_micros": 3418822, "compaction_time_cpu_micros": 785454, "output_level": 6, "num_output_files": 9, "total_output_size": 594449971, "num_input_records": 35108, "num_output_records": 34592, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 0, 0, 0, 0, 0, 9]}
The log even mentions the huge write multiplication. I wonder whether this is normal and what can be done about it.
/Z
On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk> wrote:
I need to ask here: where exactly do you observe the hundreds of GB written per day? Are the mon logs huge? Is it the mon store? Is your cluster unhealthy?
We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about 800 librbd clients. Per week our mon logs are about 70M, the cluster logs about 120M , the audit logs about 70M and I see between 100-200Kb/s writes to the mon store. That's in the lower-digit GB range per day. Hundreds of GB per day sound completely over the top on a healthy cluster, unless you have MGR modules changing the OSD/cluster map continuously.
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed by now?
@Eugen: Was there not an option to limit logging to the MON store?
For information to readers, we followed old recommendations from a Dell white paper for building a ceph cluster and have a 1TB Raid10 array on 6x write intensive SSDs for the MON stores. After 5 years we are below 10% wear. Average size of the MON store for a healthy cluster is 500M-1G, but we have seen this ballooning to 100+GB in degraded conditions.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 12:00 PM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Thank you, Eugen.
I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible.
/Z
On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
Hi,
what you report is the expected behaviour, at least I see the same on all clusters. I can't answer why the compaction is required that often, but you can control the log level of the rocksdb output:
ceph config set mon debug_rocksdb 1/5 (default is 4/5)
This reduces the log entries and you wouldn't see the manual compaction logs anymore. There are a couple more rocksdb options but I probably wouldn't change too much, only if you know what you're doing. Maybe Igor can comment if some other tuning makes sense here.
Regards, Eugen
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Any input from anyone, please?
On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
> Any input from anyone, please? > > It's another thing that seems to be rather poorly documented: it's unclear > what to expect, what 'normal' behavior should be, and what can be done > about the huge amount of writes by monitors. > > /Z > > On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <zakhar@gmail.com> wrote: > >> Hi, >> >> Monitors in our 16.2.14 cluster appear to quite often run "manual >> compaction" tasks: >> >> debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 >> {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", >> "num_memtables": 1, "num_entries": 715, "num_deletes": 251, >> "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": >> "Manual Compaction"} >> debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log >> Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] >> [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; >> will stop at (end) >> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log >> Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] >> [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; >> will stop at (end) >> debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 >> {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", >> "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, >> "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": >> "Manual Compaction"} >> debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log >> Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] >> [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; >> will stop at (end) >> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log >> Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] >> [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; >> will stop at (end) >> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 >> {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", >> "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, >> "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": >> "Manual Compaction"} >> debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log >> Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] >> [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; >> will stop at (end) >> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >> starting >> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log >> Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] >> [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; >> will stop at (end) >> >> I have removed a lot of interim log messages to save space. >> >> During each compaction the monitor process writes approximately 500-600 >> MB of data to disk over a short period of time. These writes add up to tens >> of gigabytes per hour and hundreds of gigabytes per day. >> >> Monitor rocksdb and compaction options are default: >> >> "mon_compact_on_bootstrap": "false", >> "mon_compact_on_start": "false", >> "mon_compact_on_trim": "true", >> "mon_rocksdb_options": >>
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
>> >> Is this expected behavior? Is this something I can adjust in order to >> extend the system storage life? >> >> Best regards, >> Zakhar >> > _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
9@6 files to L6, score -1.00 debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1697022610487624, "job": 70854, "event": "compaction_started", "compaction_reason": "ManualCompaction", "files_L5": [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536, 3675537, 3675538, 3675539, 3675540, 3675541], "score": -1, "input_data_size": 601117031} debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675544: 2015 keys, 67287115 bytes debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675545: 24343 keys, 67336225 bytes debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675546: 1196 keys, 67225813 bytes debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675547: 1049 keys, 67252678 bytes debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675548: 1081 keys, 67216638 bytes debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675549: 1196 keys, 67245376 bytes debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675550: 1195 keys, 67246813 bytes debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675551: 1205 keys, 67223302 bytes debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675552: 1312 keys, 56416011 bytes debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted 1@5
Thank you, Frank.
The cluster is healthy, operating normally, nothing unusual is going on. We observe lots of writes by mon processes into mon rocksdb stores, specifically:
/var/lib/ceph/mon/ceph-cephXX/store.db: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
The site of the files is not huge, but monitors rotate and write out these files often, sometimes several times per minute, resulting in lots of data written to disk. The writes coincide with "manual compaction" events logged by the monitors, for example:
debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1676] [default] [JOB 70854] Compacting 1@5
Sure, nothing unusual there: ------- cluster: id: 3f50555a-ae2a-11eb-a2fc-ffde44714d86 health: HEALTH_OK services: mon: 5 daemons, quorum ceph01,ceph03,ceph04,ceph05,ceph02 (age 2w) mgr: ceph01.vankui(active, since 12d), standbys: ceph02.shsinf osd: 96 osds: 96 up (since 2w), 95 in (since 3w) data: pools: 10 pools, 2400 pgs objects: 6.23M objects, 16 TiB usage: 61 TiB used, 716 TiB / 777 TiB avail pgs: 2396 active+clean 3 active+clean+scrubbing+deep 1 active+clean+scrubbing io: client: 2.7 GiB/s rd, 27 MiB/s wr, 46.95k op/s rd, 2.17k op/s wr ------- Please disregard the big read number, a customer is running a read-intensive job. Mon store writes keep happening when the cluster is much more quiet, thus I think that intensive reads have no effect on the mons. Mgr: "always_on_modules": [ "balancer", "crash", "devicehealth", "orchestrator", "pg_autoscaler", "progress", "rbd_support", "status", "telemetry", "volumes" ], "enabled_modules": [ "cephadm", "dashboard", "iostat", "prometheus", "restful" ], ------- /Z On Wed, 11 Oct 2023 at 14:50, Eugen Block <eblock@nde.ag> wrote:
Can you add some more details as requested by Frank? Which mgr modules are enabled? What's the current 'ceph -s' output?
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
We don't use CephFS at all and don't have RBD snapshots apart from some cloning for Openstack images.
The size of mon stores isn't an issue, it's < 600 MB. But it gets overwritten often causing lots of disk writes, and that is an issue for us.
/Z
On Wed, 11 Oct 2023 at 14:37, Eugen Block <eblock@nde.ag> wrote:
Do you use many snapshots (rbd or cephfs)? That can cause a heavy monitor usage, we've seen large mon stores on customer clusters with rbd mirroring on snapshot basis. In a healthy cluster they have mon stores of around 2GB in size.
@Eugen: Was there not an option to limit logging to the MON store?
I don't recall at the moment, worth checking tough.
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
9@6 files to L6 => 594449971 bytes debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760] [default] compacted to: base level 5 level multiplier 10.00 max bytes base 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd, 173.9 wr, level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9), read-write-amplify(3434.6) write-amplify(1707.7) OK, records in: 35108, records dropped: 516 output_compression: NoCompression debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros": 1697022613921002, "job": 70854, "event": "compaction_finished", "compaction_time_micros": 3418822, "compaction_time_cpu_micros": 785454, "output_level": 6, "num_output_files": 9, "total_output_size": 594449971, "num_input_records": 35108, "num_output_records": 34592, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 0, 0, 0, 0, 0, 9]}
The log even mentions the huge write multiplication. I wonder whether this is normal and what can be done about it.
/Z
On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk> wrote:
I need to ask here: where exactly do you observe the hundreds of GB written per day? Are the mon logs huge? Is it the mon store? Is your cluster unhealthy?
We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about 800 librbd clients. Per week our mon logs are about 70M, the cluster logs about 120M , the audit logs about 70M and I see between 100-200Kb/s writes to the mon store. That's in the lower-digit GB range per day. Hundreds of GB per day sound completely over the top on a healthy cluster, unless you have MGR modules changing the OSD/cluster map continuously.
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed by now?
@Eugen: Was there not an option to limit logging to the MON store?
For information to readers, we followed old recommendations from a Dell white paper for building a ceph cluster and have a 1TB Raid10 array on 6x write intensive SSDs for the MON stores. After 5 years we are below 10% wear. Average size of the MON store for a healthy cluster is 500M-1G, but we have seen this ballooning to 100+GB in degraded conditions.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 12:00 PM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Thank you, Eugen.
I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible.
/Z
On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
Hi,
what you report is the expected behaviour, at least I see the same on all clusters. I can't answer why the compaction is required that often, but you can control the log level of the rocksdb output:
ceph config set mon debug_rocksdb 1/5 (default is 4/5)
This reduces the log entries and you wouldn't see the manual compaction logs anymore. There are a couple more rocksdb options but I probably wouldn't change too much, only if you know what you're doing. Maybe Igor can comment if some other tuning makes sense here.
Regards, Eugen
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
> Any input from anyone, please? > > On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko < zakhar@gmail.com> wrote: > >> Any input from anyone, please? >> >> It's another thing that seems to be rather poorly documented: it's unclear >> what to expect, what 'normal' behavior should be, and what can be done >> about the huge amount of writes by monitors. >> >> /Z >> >> On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko < zakhar@gmail.com> wrote: >> >>> Hi, >>> >>> Monitors in our 16.2.14 cluster appear to quite often run "manual >>> compaction" tasks: >>> >>> debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 >>> {"time_micros": 1696843853892760, "job": 64225, "event": "flush_started", >>> "num_memtables": 1, "num_entries": 715, "num_deletes": 251, >>> "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": >>> "Manual Compaction"} >>> debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log >>> Time 2023/10/09-09:30:53.910204) [db_impl/db_impl_compaction_flush.cc:2516] >>> [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; >>> will stop at (end) >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original Log >>> Time 2023/10/09-09:30:53.911004) [db_impl/db_impl_compaction_flush.cc:2516] >>> [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; >>> will stop at (end) >>> debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 >>> {"time_micros": 1696843928961390, "job": 64228, "event": "flush_started", >>> "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, >>> "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": >>> "Manual Compaction"} >>> debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log >>> Time 2023/10/09-09:32:08.977739) [db_impl/db_impl_compaction_flush.cc:2516] >>> [default] Manual compaction from level-0 to level-5 from 'logm .. 'logm; >>> will stop at (end) >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original Log >>> Time 2023/10/09-09:32:08.978512) [db_impl/db_impl_compaction_flush.cc:2516] >>> [default] Manual compaction from level-5 to level-6 from 'logm .. 'logm; >>> will stop at (end) >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: EVENT_LOG_v1 >>> {"time_micros": 1696844009033151, "job": 64231, "event": "flush_started", >>> "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, >>> "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": >>> "Manual Compaction"} >>> debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log >>> Time 2023/10/09-09:33:29.049585) [db_impl/db_impl_compaction_flush.cc:2516] >>> [default] Manual compaction from level-0 to level-5 from 'paxos .. 'paxos; >>> will stop at (end) >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction >>> starting >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original Log >>> Time 2023/10/09-09:33:29.050355) [db_impl/db_impl_compaction_flush.cc:2516] >>> [default] Manual compaction from level-5 to level-6 from 'paxos .. 'paxos; >>> will stop at (end) >>> >>> I have removed a lot of interim log messages to save space. >>> >>> During each compaction the monitor process writes approximately 500-600 >>> MB of data to disk over a short period of time. These writes add up to tens >>> of gigabytes per hour and hundreds of gigabytes per day. >>> >>> Monitor rocksdb and compaction options are default: >>> >>> "mon_compact_on_bootstrap": "false", >>> "mon_compact_on_start": "false", >>> "mon_compact_on_trim": "true", >>> "mon_rocksdb_options": >>>
9@6 files to L6, score -1.00 debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1697022610487624, "job": 70854, "event": "compaction_started", "compaction_reason": "ManualCompaction", "files_L5": [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536, 3675537, 3675538, 3675539, 3675540, 3675541], "score": -1, "input_data_size": 601117031} debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675544: 2015 keys, 67287115 bytes debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675545: 24343 keys, 67336225 bytes debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675546: 1196 keys, 67225813 bytes debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675547: 1049 keys, 67252678 bytes debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675548: 1081 keys, 67216638 bytes debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675549: 1196 keys, 67245376 bytes debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675550: 1195 keys, 67246813 bytes debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675551: 1205 keys, 67223302 bytes debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675552: 1312 keys, 56416011 bytes debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted 1@5
Thank you, Frank.
The cluster is healthy, operating normally, nothing unusual is going on. We observe lots of writes by mon processes into mon rocksdb stores, specifically:
/var/lib/ceph/mon/ceph-cephXX/store.db: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
The site of the files is not huge, but monitors rotate and write out these files often, sometimes several times per minute, resulting in lots of data written to disk. The writes coincide with "manual compaction" events logged by the monitors, for example:
debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1676] [default] [JOB 70854] Compacting 1@5
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
>>> >>> Is this expected behavior? Is this something I can adjust in order to >>> extend the system storage life? >>> >>> Best regards, >>> Zakhar >>> >> > _______________________________________________ > ceph-users mailing list -- ceph-users@ceph.io > To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
That all looks normal to me, to be honest. Can you show some details how you calculate the "hundreds of GB per day"? I see similar stats as Frank on different clusters with different client IO. Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Sure, nothing unusual there:
-------
cluster: id: 3f50555a-ae2a-11eb-a2fc-ffde44714d86 health: HEALTH_OK
services: mon: 5 daemons, quorum ceph01,ceph03,ceph04,ceph05,ceph02 (age 2w) mgr: ceph01.vankui(active, since 12d), standbys: ceph02.shsinf osd: 96 osds: 96 up (since 2w), 95 in (since 3w)
data: pools: 10 pools, 2400 pgs objects: 6.23M objects, 16 TiB usage: 61 TiB used, 716 TiB / 777 TiB avail pgs: 2396 active+clean 3 active+clean+scrubbing+deep 1 active+clean+scrubbing
io: client: 2.7 GiB/s rd, 27 MiB/s wr, 46.95k op/s rd, 2.17k op/s wr
-------
Please disregard the big read number, a customer is running a read-intensive job. Mon store writes keep happening when the cluster is much more quiet, thus I think that intensive reads have no effect on the mons.
Mgr:
"always_on_modules": [ "balancer", "crash", "devicehealth", "orchestrator", "pg_autoscaler", "progress", "rbd_support", "status", "telemetry", "volumes" ], "enabled_modules": [ "cephadm", "dashboard", "iostat", "prometheus", "restful" ],
-------
/Z
On Wed, 11 Oct 2023 at 14:50, Eugen Block <eblock@nde.ag> wrote:
Can you add some more details as requested by Frank? Which mgr modules are enabled? What's the current 'ceph -s' output?
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
We don't use CephFS at all and don't have RBD snapshots apart from some cloning for Openstack images.
The size of mon stores isn't an issue, it's < 600 MB. But it gets overwritten often causing lots of disk writes, and that is an issue for us.
/Z
On Wed, 11 Oct 2023 at 14:37, Eugen Block <eblock@nde.ag> wrote:
Do you use many snapshots (rbd or cephfs)? That can cause a heavy monitor usage, we've seen large mon stores on customer clusters with rbd mirroring on snapshot basis. In a healthy cluster they have mon stores of around 2GB in size.
@Eugen: Was there not an option to limit logging to the MON store?
I don't recall at the moment, worth checking tough.
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
9@6 files to L6 => 594449971 bytes debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760] [default] compacted to: base level 5 level multiplier 10.00 max bytes base 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd, 173.9 wr, level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9), read-write-amplify(3434.6) write-amplify(1707.7) OK, records in: 35108, records dropped: 516 output_compression: NoCompression debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros": 1697022613921002, "job": 70854, "event": "compaction_finished", "compaction_time_micros": 3418822, "compaction_time_cpu_micros": 785454, "output_level": 6, "num_output_files": 9, "total_output_size": 594449971, "num_input_records": 35108, "num_output_records": 34592, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 0, 0, 0, 0, 0, 9]}
The log even mentions the huge write multiplication. I wonder whether this is normal and what can be done about it.
/Z
On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk> wrote:
I need to ask here: where exactly do you observe the hundreds of GB written per day? Are the mon logs huge? Is it the mon store? Is your cluster unhealthy?
We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about 800 librbd clients. Per week our mon logs are about 70M, the cluster logs about 120M , the audit logs about 70M and I see between 100-200Kb/s writes to the mon store. That's in the lower-digit GB range per day. Hundreds of GB per day sound completely over the top on a healthy cluster, unless you have MGR modules changing the OSD/cluster map continuously.
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed by now?
@Eugen: Was there not an option to limit logging to the MON store?
For information to readers, we followed old recommendations from a Dell white paper for building a ceph cluster and have a 1TB Raid10 array on 6x write intensive SSDs for the MON stores. After 5 years we are below 10% wear. Average size of the MON store for a healthy cluster is 500M-1G, but we have seen this ballooning to 100+GB in degraded conditions.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 12:00 PM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Thank you, Eugen.
I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible.
/Z
On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
> Hi, > > what you report is the expected behaviour, at least I see the same on > all clusters. I can't answer why the compaction is required that > often, but you can control the log level of the rocksdb output: > > ceph config set mon debug_rocksdb 1/5 (default is 4/5) > > This reduces the log entries and you wouldn't see the manual > compaction logs anymore. There are a couple more rocksdb options but I > probably wouldn't change too much, only if you know what you're doing. > Maybe Igor can comment if some other tuning makes sense here. > > Regards, > Eugen > > Zitat von Zakhar Kirpichenko <zakhar@gmail.com>: > > > Any input from anyone, please? > > > > On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko < zakhar@gmail.com> > wrote: > > > >> Any input from anyone, please? > >> > >> It's another thing that seems to be rather poorly documented: it's > unclear > >> what to expect, what 'normal' behavior should be, and what can be done > >> about the huge amount of writes by monitors. > >> > >> /Z > >> > >> On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko < zakhar@gmail.com> > wrote: > >> > >>> Hi, > >>> > >>> Monitors in our 16.2.14 cluster appear to quite often run "manual > >>> compaction" tasks: > >>> > >>> debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: > EVENT_LOG_v1 > >>> {"time_micros": 1696843853892760, "job": 64225, "event": > "flush_started", > >>> "num_memtables": 1, "num_entries": 715, "num_deletes": 251, > >>> "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": > >>> "Manual Compaction"} > >>> debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:30:53.910204) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-0 to level-5 from 'paxos .. > 'paxos; > >>> will stop at (end) > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:30:53.911004) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-5 to level-6 from 'paxos .. > 'paxos; > >>> will stop at (end) > >>> debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: > EVENT_LOG_v1 > >>> {"time_micros": 1696843928961390, "job": 64228, "event": > "flush_started", > >>> "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, > >>> "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": > >>> "Manual Compaction"} > >>> debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:32:08.977739) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-0 to level-5 from 'logm .. > 'logm; > >>> will stop at (end) > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:32:08.978512) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-5 to level-6 from 'logm .. > 'logm; > >>> will stop at (end) > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: > EVENT_LOG_v1 > >>> {"time_micros": 1696844009033151, "job": 64231, "event": > "flush_started", > >>> "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, > >>> "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": > >>> "Manual Compaction"} > >>> debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:33:29.049585) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-0 to level-5 from 'paxos .. > 'paxos; > >>> will stop at (end) > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:33:29.050355) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-5 to level-6 from 'paxos .. > 'paxos; > >>> will stop at (end) > >>> > >>> I have removed a lot of interim log messages to save space. > >>> > >>> During each compaction the monitor process writes approximately 500-600 > >>> MB of data to disk over a short period of time. These writes add up to > tens > >>> of gigabytes per hour and hundreds of gigabytes per day. > >>> > >>> Monitor rocksdb and compaction options are default: > >>> > >>> "mon_compact_on_bootstrap": "false", > >>> "mon_compact_on_start": "false", > >>> "mon_compact_on_trim": "true", > >>> "mon_rocksdb_options": > >>> >
9@6 files to L6, score -1.00 debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1697022610487624, "job": 70854, "event": "compaction_started", "compaction_reason": "ManualCompaction", "files_L5": [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536, 3675537, 3675538, 3675539, 3675540, 3675541], "score": -1, "input_data_size": 601117031} debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675544: 2015 keys, 67287115 bytes debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675545: 24343 keys, 67336225 bytes debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675546: 1196 keys, 67225813 bytes debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675547: 1049 keys, 67252678 bytes debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675548: 1081 keys, 67216638 bytes debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675549: 1196 keys, 67245376 bytes debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675550: 1195 keys, 67246813 bytes debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675551: 1205 keys, 67223302 bytes debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675552: 1312 keys, 56416011 bytes debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted 1@5
Thank you, Frank.
The cluster is healthy, operating normally, nothing unusual is going on. We observe lots of writes by mon processes into mon rocksdb stores, specifically:
/var/lib/ceph/mon/ceph-cephXX/store.db: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
The site of the files is not huge, but monitors rotate and write out these files often, sometimes several times per minute, resulting in lots of data written to disk. The writes coincide with "manual compaction" events logged by the monitors, for example:
debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1676] [default] [JOB 70854] Compacting 1@5
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
> >>> > >>> Is this expected behavior? Is this something I can adjust in order to > >>> extend the system storage life? > >>> > >>> Best regards, > >>> Zakhar > >>> > >> > > _______________________________________________ > > ceph-users mailing list -- ceph-users@ceph.io > > To unsubscribe send an email to ceph-users-leave@ceph.io > > > _______________________________________________ > ceph-users mailing list -- ceph-users@ceph.io > To unsubscribe send an email to ceph-users-leave@ceph.io > _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Eugen,
Thanks for your response. May I ask what numbers you're referring to?
I am not referring to monitor store.db sizes. I am specifically referring
to writes monitors do to their store.db file by frequently rotating and
replacing them with new versions during compactions. The size of the
store.db remains more or less the same.
This is a 300s iotop snippet, sorted by aggregated disk writes:
Total DISK READ: 35.56 M/s | Total DISK WRITE: 23.89 M/s
Current DISK READ: 35.64 M/s | Current DISK WRITE: 24.09 M/s
TID PRIO USER DISK READ DISK WRITE> SWAPIN IO COMMAND
4919 be/4 167 16.75 M 2.24 G 0.00 % 1.34 % ceph-mon -n
mon.ceph03 -f --setuser ceph --setgr~lt-mon-cluster-log-to-stderr=true
[rocksdb:low0]
15122 be/4 167 0.00 B 652.91 M 0.00 % 0.27 % ceph-osd -n
osd.31 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
17073 be/4 167 0.00 B 651.86 M 0.00 % 0.27 % ceph-osd -n
osd.32 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
17268 be/4 167 0.00 B 490.86 M 0.00 % 0.18 % ceph-osd -n
osd.25 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
18032 be/4 167 0.00 B 463.57 M 0.00 % 0.17 % ceph-osd -n
osd.26 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
16855 be/4 167 0.00 B 402.86 M 0.00 % 0.15 % ceph-osd -n
osd.22 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
17406 be/4 167 0.00 B 387.03 M 0.00 % 0.14 % ceph-osd -n
osd.27 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
17932 be/4 167 0.00 B 375.42 M 0.00 % 0.13 % ceph-osd -n
osd.29 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
18017 be/4 167 0.00 B 359.38 M 0.00 % 0.13 % ceph-osd -n
osd.28 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
17420 be/4 167 0.00 B 332.83 M 0.00 % 0.12 % ceph-osd -n
osd.23 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
17975 be/4 167 0.00 B 312.06 M 0.00 % 0.11 % ceph-osd -n
osd.30 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
17273 be/4 167 0.00 B 303.49 M 0.00 % 0.11 % ceph-osd -n
osd.24 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
[bstore_kv_sync]
Not a good example, because sometimes mon writes more intensively, but it
is very apparent that thread 4919 of the monitor process is the top disk
writer in the system.
This is the mon thread producing lots of writes:
4919 167 20 0 2031116 1.1g 10652 S 0.0 0.3 288:48.65
rocksdb:low0
Then with a combination of lsof and sysdig I determine that the writes are
being made to /var/lib/ceph/mon/ceph-ceph03/store.db/*.sst, i.e. the mon's
rocksdb store:
ceph-mon 4838 167 200r REG 253,11 67319253 14812899
/var/lib/ceph/mon/ceph-ceph03/store.db/3677146.sst
ceph-mon 4838 167 203r REG 253,11 67228736 14813270
/var/lib/ceph/mon/ceph-ceph03/store.db/3677147.sst
ceph-mon 4838 167 205r REG 253,11 67243212 14813275
/var/lib/ceph/mon/ceph-ceph03/store.db/3677148.sst
ceph-mon 4838 167 208r REG 253,11 67247953 14813316
/var/lib/ceph/mon/ceph-ceph03/store.db/3677149.sst
ceph-mon 4838 167 220r REG 253,11 67261659 14813332
/var/lib/ceph/mon/ceph-ceph03/store.db/3677150.sst
ceph-mon 4838 167 221r REG 253,11 67242500 14813345
/var/lib/ceph/mon/ceph-ceph03/store.db/3677151.sst
ceph-mon 4838 167 224r REG 253,11 67264969 14813348
/var/lib/ceph/mon/ceph-ceph03/store.db/3677152.sst
ceph-mon 4838 167 228r REG 253,11 64346933 14813381
/var/lib/ceph/mon/ceph-ceph03/store.db/3677153.sst
By matching iotop and sysdig write records to mon's log entries, I see that
the writes happen during "manual compaction" events - whatever they are,
because there's no documentation on this whatsoever, and each time around
0.56GB is being written to disk to a new set of *.sst files, which is the
total size of the store.db. Looks like from time to time the monitor just
reads its store.db and writes it out to a new set of files, as the file
names "numbers" increase with each write:
ceph-mon 4838 167 175r REG 253,11 67220863 14812310
/var/lib/ceph/mon/ceph-ceph03/store.db/3677167.sst
ceph-mon 4838 167 200r REG 253,11 67358627 14812899
/var/lib/ceph/mon/ceph-ceph03/store.db/3677168.sst
ceph-mon 4838 167 203r REG 253,11 67277978 14813270
/var/lib/ceph/mon/ceph-ceph03/store.db/3677169.sst
ceph-mon 4838 167 205r REG 253,11 67256312 14813275
/var/lib/ceph/mon/ceph-ceph03/store.db/3677170.sst
ceph-mon 4838 167 208r REG 253,11 67226761 14813316
/var/lib/ceph/mon/ceph-ceph03/store.db/3677171.sst
ceph-mon 4838 167 220r REG 253,11 67258798 14813332
/var/lib/ceph/mon/ceph-ceph03/store.db/3677172.sst
ceph-mon 4838 167 221r REG 253,11 67224665 14813345
/var/lib/ceph/mon/ceph-ceph03/store.db/3677173.sst
ceph-mon 4838 167 224r REG 253,11 67224123 14813348
/var/lib/ceph/mon/ceph-ceph03/store.db/3677174.sst
ceph-mon 4838 167 228r REG 253,11 62195349 14813381
/var/lib/ceph/mon/ceph-ceph03/store.db/3677175.sst
I hope this clears up the situation.
Do you observe this behavior in your clusters? Can you please check whether
your mons do something similar and store.db/*.sst change often?
/Z
On Wed, 11 Oct 2023 at 16:22, Eugen Block <eblock@nde.ag> wrote:
> That all looks normal to me, to be honest. Can you show some details
> how you calculate the "hundreds of GB per day"? I see similar stats as
> Frank on different clusters with different client IO.
>
> Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
>
> > Sure, nothing unusual there:
> >
> > -------
> >
> > cluster:
> > id: 3f50555a-ae2a-11eb-a2fc-ffde44714d86
> > health: HEALTH_OK
> >
> > services:
> > mon: 5 daemons, quorum ceph01,ceph03,ceph04,ceph05,ceph02 (age 2w)
> > mgr: ceph01.vankui(active, since 12d), standbys: ceph02.shsinf
> > osd: 96 osds: 96 up (since 2w), 95 in (since 3w)
> >
> > data:
> > pools: 10 pools, 2400 pgs
> > objects: 6.23M objects, 16 TiB
> > usage: 61 TiB used, 716 TiB / 777 TiB avail
> > pgs: 2396 active+clean
> > 3 active+clean+scrubbing+deep
> > 1 active+clean+scrubbing
> >
> > io:
> > client: 2.7 GiB/s rd, 27 MiB/s wr, 46.95k op/s rd, 2.17k op/s wr
> >
> > -------
> >
> > Please disregard the big read number, a customer is running a
> > read-intensive job. Mon store writes keep happening when the cluster is
> > much more quiet, thus I think that intensive reads have no effect on the
> > mons.
> >
> > Mgr:
> >
> > "always_on_modules": [
> > "balancer",
> > "crash",
> > "devicehealth",
> > "orchestrator",
> > "pg_autoscaler",
> > "progress",
> > "rbd_support",
> > "status",
> > "telemetry",
> > "volumes"
> > ],
> > "enabled_modules": [
> > "cephadm",
> > "dashboard",
> > "iostat",
> > "prometheus",
> > "restful"
> > ],
> >
> > -------
> >
> > /Z
> >
> >
> > On Wed, 11 Oct 2023 at 14:50, Eugen Block <eblock@nde.ag> wrote:
> >
> >> Can you add some more details as requested by Frank? Which mgr modules
> >> are enabled? What's the current 'ceph -s' output?
> >>
> >> > Is autoscaler running and doing stuff?
> >> > Is balancer running and doing stuff?
> >> > Is backfill going on?
> >> > Is recovery going on?
> >> > Is your ceph version affected by the "excessive logging to MON
> >> > store" issue that was present starting with pacific but should have
> >> > been addressed
> >>
> >>
> >> Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
> >>
> >> > We don't use CephFS at all and don't have RBD snapshots apart from
> some
> >> > cloning for Openstack images.
> >> >
> >> > The size of mon stores isn't an issue, it's < 600 MB. But it gets
> >> > overwritten often causing lots of disk writes, and that is an issue
> for
> >> us.
> >> >
> >> > /Z
> >> >
> >> > On Wed, 11 Oct 2023 at 14:37, Eugen Block <eblock@nde.ag> wrote:
> >> >
> >> >> Do you use many snapshots (rbd or cephfs)? That can cause a heavy
> >> >> monitor usage, we've seen large mon stores on customer clusters with
> >> >> rbd mirroring on snapshot basis. In a healthy cluster they have mon
> >> >> stores of around 2GB in size.
> >> >>
> >> >> >> @Eugen: Was there not an option to limit logging to the MON store?
> >> >>
> >> >> I don't recall at the moment, worth checking tough.
> >> >>
> >> >> Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
> >> >>
> >> >> > Thank you, Frank.
> >> >> >
> >> >> > The cluster is healthy, operating normally, nothing unusual is
> going
> >> on.
> >> >> We
> >> >> > observe lots of writes by mon processes into mon rocksdb stores,
> >> >> > specifically:
> >> >> >
> >> >> > /var/lib/ceph/mon/ceph-cephXX/store.db:
> >> >> > 65M 3675511.sst
> >> >> > 65M 3675512.sst
> >> >> > 65M 3675513.sst
> >> >> > 65M 3675514.sst
> >> >> > 65M 3675515.sst
> >> >> > 65M 3675516.sst
> >> >> > 65M 3675517.sst
> >> >> > 65M 3675518.sst
> >> >> > 62M 3675519.sst
> >> >> >
> >> >> > The site of the files is not huge, but monitors rotate and write
> out
> >> >> these
> >> >> > files often, sometimes several times per minute, resulting in lots
> of
> >> >> data
> >> >> > written to disk. The writes coincide with "manual compaction"
> events
> >> >> logged
> >> >> > by the monitors, for example:
> >> >> >
> >> >> > debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1676] [default] [JOB 70854]
> Compacting
> >> 1@5
> >> >> +
> >> >> > 9@6 files to L6, score -1.00
> >> >> > debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb:
> >> EVENT_LOG_v1
> >> >> > {"time_micros": 1697022610487624, "job": 70854, "event":
> >> >> > "compaction_started", "compaction_reason": "ManualCompaction",
> >> >> "files_L5":
> >> >> > [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536,
> 3675537,
> >> >> > 3675538, 3675539, 3675540, 3675541], "score": -1,
> "input_data_size":
> >> >> > 601117031}
> >> >> > debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675544: 2015 keys, 67287115 bytes
> >> >> > debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675545: 24343 keys, 67336225 bytes
> >> >> > debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675546: 1196 keys, 67225813 bytes
> >> >> > debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675547: 1049 keys, 67252678 bytes
> >> >> > debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675548: 1081 keys, 67216638 bytes
> >> >> > debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675549: 1196 keys, 67245376 bytes
> >> >> > debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675550: 1195 keys, 67246813 bytes
> >> >> > debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675551: 1205 keys, 67223302 bytes
> >> >> > debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675552: 1312 keys, 56416011 bytes
> >> >> > debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted
> >> 1@5
> >> >> +
> >> >> > 9@6 files to L6 => 594449971 bytes
> >> >> > debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb:
> (Original
> >> Log
> >> >> > Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760]
> >> >> > [default] compacted to: base level 5 level multiplier 10.00 max
> bytes
> >> >> base
> >> >> > 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd,
> 173.9
> >> >> wr,
> >> >> > level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9),
> >> >> > read-write-amplify(3434.6) write-amplify(1707.7) OK, records in:
> >> 35108,
> >> >> > records dropped: 516 output_compression: NoCompression
> >> >> > debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb:
> (Original
> >> Log
> >> >> > Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros":
> >> >> > 1697022613921002, "job": 70854, "event": "compaction_finished",
> >> >> > "compaction_time_micros": 3418822, "compaction_time_cpu_micros":
> >> 785454,
> >> >> > "output_level": 6, "num_output_files": 9, "total_output_size":
> >> 594449971,
> >> >> > "num_input_records": 35108, "num_output_records": 34592,
> >> >> > "num_subcompactions": 1, "output_compression": "NoCompression",
> >> >> > "num_single_delete_mismatches": 0,
> "num_single_delete_fallthrough": 0,
> >> >> > "lsm_state": [0, 0, 0, 0, 0, 0, 9]}
> >> >> >
> >> >> > The log even mentions the huge write multiplication. I wonder
> whether
> >> >> this
> >> >> > is normal and what can be done about it.
> >> >> >
> >> >> > /Z
> >> >> >
> >> >> > On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk> wrote:
> >> >> >
> >> >> >> I need to ask here: where exactly do you observe the hundreds of
> GB
> >> >> >> written per day? Are the mon logs huge? Is it the mon store? Is
> your
> >> >> >> cluster unhealthy?
> >> >> >>
> >> >> >> We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients
> and
> >> >> about
> >> >> >> 800 librbd clients. Per week our mon logs are about 70M, the
> cluster
> >> >> logs
> >> >> >> about 120M , the audit logs about 70M and I see between
> 100-200Kb/s
> >> >> writes
> >> >> >> to the mon store. That's in the lower-digit GB range per day.
> >> Hundreds
> >> >> of
> >> >> >> GB per day sound completely over the top on a healthy cluster,
> unless
> >> >> you
> >> >> >> have MGR modules changing the OSD/cluster map continuously.
> >> >> >>
> >> >> >> Is autoscaler running and doing stuff?
> >> >> >> Is balancer running and doing stuff?
> >> >> >> Is backfill going on?
> >> >> >> Is recovery going on?
> >> >> >> Is your ceph version affected by the "excessive logging to MON
> store"
> >> >> >> issue that was present starting with pacific but should have been
> >> >> addressed
> >> >> >> by now?
> >> >> >>
> >> >> >> @Eugen: Was there not an option to limit logging to the MON store?
> >> >> >>
> >> >> >> For information to readers, we followed old recommendations from a
> >> Dell
> >> >> >> white paper for building a ceph cluster and have a 1TB Raid10
> array
> >> on
> >> >> 6x
> >> >> >> write intensive SSDs for the MON stores. After 5 years we are
> below
> >> 10%
> >> >> >> wear. Average size of the MON store for a healthy cluster is
> 500M-1G,
> >> >> but
> >> >> >> we have seen this ballooning to 100+GB in degraded conditions.
> >> >> >>
> >> >> >> Best regards,
> >> >> >> =================
> >> >> >> Frank Schilder
> >> >> >> AIT Risø Campus
> >> >> >> Bygning 109, rum S14
> >> >> >>
> >> >> >> ________________________________________
> >> >> >> From: Zakhar Kirpichenko <zakhar@gmail.com>
> >> >> >> Sent: Wednesday, October 11, 2023 12:00 PM
> >> >> >> To: Eugen Block
> >> >> >> Cc: ceph-users@ceph.io
> >> >> >> Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
> >> >> >>
> >> >> >> Thank you, Eugen.
> >> >> >>
> >> >> >> I'm interested specifically to find out whether the huge amount of
> >> data
> >> >> >> written by monitors is expected. It is eating through the
> endurance
> >> of
> >> >> our
> >> >> >> system drives, which were not specced for high DWPD/TBW, as this
> is
> >> not
> >> >> a
> >> >> >> documented requirement, and monitors produce hundreds of
> gigabytes of
> >> >> >> writes per day. I am looking for ways to reduce the amount of
> >> writes, if
> >> >> >> possible.
> >> >> >>
> >> >> >> /Z
> >> >> >>
> >> >> >> On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag> wrote:
> >> >> >>
> >> >> >> > Hi,
> >> >> >> >
> >> >> >> > what you report is the expected behaviour, at least I see the
> same
> >> on
> >> >> >> > all clusters. I can't answer why the compaction is required that
> >> >> >> > often, but you can control the log level of the rocksdb output:
> >> >> >> >
> >> >> >> > ceph config set mon debug_rocksdb 1/5 (default is 4/5)
> >> >> >> >
> >> >> >> > This reduces the log entries and you wouldn't see the manual
> >> >> >> > compaction logs anymore. There are a couple more rocksdb options
> >> but I
> >> >> >> > probably wouldn't change too much, only if you know what you're
> >> doing.
> >> >> >> > Maybe Igor can comment if some other tuning makes sense here.
> >> >> >> >
> >> >> >> > Regards,
> >> >> >> > Eugen
> >> >> >> >
> >> >> >> > Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
> >> >> >> >
> >> >> >> > > Any input from anyone, please?
> >> >> >> > >
> >> >> >> > > On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <
> >> zakhar@gmail.com>
> >> >> >> > wrote:
> >> >> >> > >
> >> >> >> > >> Any input from anyone, please?
> >> >> >> > >>
> >> >> >> > >> It's another thing that seems to be rather poorly documented:
> >> it's
> >> >> >> > unclear
> >> >> >> > >> what to expect, what 'normal' behavior should be, and what
> can
> >> be
> >> >> done
> >> >> >> > >> about the huge amount of writes by monitors.
> >> >> >> > >>
> >> >> >> > >> /Z
> >> >> >> > >>
> >> >> >> > >> On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <
> >> zakhar@gmail.com>
> >> >> >> > wrote:
> >> >> >> > >>
> >> >> >> > >>> Hi,
> >> >> >> > >>>
> >> >> >> > >>> Monitors in our 16.2.14 cluster appear to quite often run
> >> "manual
> >> >> >> > >>> compaction" tasks:
> >> >> >> > >>>
> >> >> >> > >>> debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb:
> >> >> >> > EVENT_LOG_v1
> >> >> >> > >>> {"time_micros": 1696843853892760, "job": 64225, "event":
> >> >> >> > "flush_started",
> >> >> >> > >>> "num_memtables": 1, "num_entries": 715, "num_deletes": 251,
> >> >> >> > >>> "total_data_size": 3870352, "memory_usage": 3886744,
> >> >> "flush_reason":
> >> >> >> > >>> "Manual Compaction"}
> >> >> >> > >>> debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:30:53.910204)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-0 to level-5 from
> >> 'paxos ..
> >> >> >> > 'paxos;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:30:53.911004)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-5 to level-6 from
> >> 'paxos ..
> >> >> >> > 'paxos;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb:
> >> >> >> > EVENT_LOG_v1
> >> >> >> > >>> {"time_micros": 1696843928961390, "job": 64228, "event":
> >> >> >> > "flush_started",
> >> >> >> > >>> "num_memtables": 1, "num_entries": 1580, "num_deletes": 502,
> >> >> >> > >>> "total_data_size": 8404605, "memory_usage": 8465840,
> >> >> "flush_reason":
> >> >> >> > >>> "Manual Compaction"}
> >> >> >> > >>> debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:32:08.977739)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-0 to level-5 from
> 'logm
> >> ..
> >> >> >> > 'logm;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:32:08.978512)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-5 to level-6 from
> 'logm
> >> ..
> >> >> >> > 'logm;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb:
> >> >> >> > EVENT_LOG_v1
> >> >> >> > >>> {"time_micros": 1696844009033151, "job": 64231, "event":
> >> >> >> > "flush_started",
> >> >> >> > >>> "num_memtables": 1, "num_entries": 1430, "num_deletes": 251,
> >> >> >> > >>> "total_data_size": 8975535, "memory_usage": 9035920,
> >> >> "flush_reason":
> >> >> >> > >>> "Manual Compaction"}
> >> >> >> > >>> debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:33:29.049585)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-0 to level-5 from
> >> 'paxos ..
> >> >> >> > 'paxos;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:33:29.050355)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-5 to level-6 from
> >> 'paxos ..
> >> >> >> > 'paxos;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>>
> >> >> >> > >>> I have removed a lot of interim log messages to save space.
> >> >> >> > >>>
> >> >> >> > >>> During each compaction the monitor process writes
> approximately
> >> >> >> 500-600
> >> >> >> > >>> MB of data to disk over a short period of time. These writes
> >> add
> >> >> up
> >> >> >> to
> >> >> >> > tens
> >> >> >> > >>> of gigabytes per hour and hundreds of gigabytes per day.
> >> >> >> > >>>
> >> >> >> > >>> Monitor rocksdb and compaction options are default:
> >> >> >> > >>>
> >> >> >> > >>> "mon_compact_on_bootstrap": "false",
> >> >> >> > >>> "mon_compact_on_start": "false",
> >> >> >> > >>> "mon_compact_on_trim": "true",
> >> >> >> > >>> "mon_rocksdb_options":
> >> >> >> > >>>
> >> >> >> >
> >> >> >>
> >> >>
> >>
> "write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
> >> >> >> > >>>
> >> >> >> > >>> Is this expected behavior? Is this something I can adjust in
> >> >> order to
> >> >> >> > >>> extend the system storage life?
> >> >> >> > >>>
> >> >> >> > >>> Best regards,
> >> >> >> > >>> Zakhar
> >> >> >> > >>>
> >> >> >> > >>
> >> >> >> > > _______________________________________________
> >> >> >> > > ceph-users mailing list -- ceph-users@ceph.io
> >> >> >> > > To unsubscribe send an email to ceph-users-leave@ceph.io
> >> >> >> >
> >> >> >> >
> >> >> >> > _______________________________________________
> >> >> >> > ceph-users mailing list -- ceph-users@ceph.io
> >> >> >> > To unsubscribe send an email to ceph-users-leave@ceph.io
> >> >> >> >
> >> >> >> _______________________________________________
> >> >> >> ceph-users mailing list -- ceph-users@ceph.io
> >> >> >> To unsubscribe send an email to ceph-users-leave@ceph.io
> >> >> >>
> >> >>
> >> >>
> >> >>
> >> >>
> >>
> >>
> >>
> >>
>
>
>
>
Oh wow! I never bothered looking, because on our hardware the wear is so low: # iotop -ao -bn 2 -d 300 Total DISK READ : 0.00 B/s | Total DISK WRITE : 6.46 M/s Actual DISK READ: 0.00 B/s | Actual DISK WRITE: 6.47 M/s TID PRIO USER DISK READ DISK WRITE SWAPIN IO COMMAND 2230 be/4 ceph 0.00 B 1818.71 M 0.00 % 0.46 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [rocksdb:low0] 2256 be/4 ceph 0.00 B 19.27 M 0.00 % 0.43 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [safe_timer] 2250 be/4 ceph 0.00 B 42.38 M 0.00 % 0.26 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [fn_monstore] 2231 be/4 ceph 0.00 B 58.36 M 0.00 % 0.01 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [rocksdb:high0] 644 be/3 root 0.00 B 576.00 K 0.00 % 0.00 % [jbd2/sda3-8] 2225 be/4 ceph 0.00 B 128.00 K 0.00 % 0.00 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [log] 1637141 be/4 root 0.00 B 0.00 B 0.00 % 0.00 % [kworker/u113:2-flush-8:0] 1636453 be/4 root 0.00 B 0.00 B 0.00 % 0.00 % [kworker/u112:0-ceph0] 1560 be/4 root 0.00 B 20.00 K 0.00 % 0.00 % rsyslogd -n [in:imjournal] 1561 be/4 root 0.00 B 56.00 K 0.00 % 0.00 % rsyslogd -n [rs:main Q:Reg] 1.8GB every 5 minutes, thats 518GB per day. The 400G drives we have are rated 10DWPD and with the 6-drives RAID10 config this gives plenty of life-time. I guess this write load will kill any low-grade SSD (typical bood devices, even enterprise ones) specifically if its smaller drives and the controller doesn't reallocate cells according to remaining write endurance. I guess there was a reason for the recommendations by Dell. I always thought that the recent recommendation for MON store storage in the ceph docs are a "bit unrealistic", apparently both, in size and in performance (including endurance). Well, I guess you need to look for write intensive drives with decent specs. If you do, also go for sufficient size. This will absorb temporary usage peaks that can be very large and also provide extra endurance with SSDs with good controllers. I also think the recommendations on the ceph docs deserve a reality check. Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14 ________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 4:30 PM To: Eugen Block Cc: Frank Schilder; ceph-users@ceph.io Subject: Re: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes Eugen, Thanks for your response. May I ask what numbers you're referring to? I am not referring to monitor store.db sizes. I am specifically referring to writes monitors do to their store.db file by frequently rotating and replacing them with new versions during compactions. The size of the store.db remains more or less the same. This is a 300s iotop snippet, sorted by aggregated disk writes: Total DISK READ: 35.56 M/s | Total DISK WRITE: 23.89 M/s Current DISK READ: 35.64 M/s | Current DISK WRITE: 24.09 M/s TID PRIO USER DISK READ DISK WRITE> SWAPIN IO COMMAND 4919 be/4 167 16.75 M 2.24 G 0.00 % 1.34 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgr~lt-mon-cluster-log-to-stderr=true [rocksdb:low0] 15122 be/4 167 0.00 B 652.91 M 0.00 % 0.27 % ceph-osd -n osd.31 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17073 be/4 167 0.00 B 651.86 M 0.00 % 0.27 % ceph-osd -n osd.32 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17268 be/4 167 0.00 B 490.86 M 0.00 % 0.18 % ceph-osd -n osd.25 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 18032 be/4 167 0.00 B 463.57 M 0.00 % 0.17 % ceph-osd -n osd.26 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 16855 be/4 167 0.00 B 402.86 M 0.00 % 0.15 % ceph-osd -n osd.22 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17406 be/4 167 0.00 B 387.03 M 0.00 % 0.14 % ceph-osd -n osd.27 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17932 be/4 167 0.00 B 375.42 M 0.00 % 0.13 % ceph-osd -n osd.29 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 18017 be/4 167 0.00 B 359.38 M 0.00 % 0.13 % ceph-osd -n osd.28 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17420 be/4 167 0.00 B 332.83 M 0.00 % 0.12 % ceph-osd -n osd.23 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17975 be/4 167 0.00 B 312.06 M 0.00 % 0.11 % ceph-osd -n osd.30 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17273 be/4 167 0.00 B 303.49 M 0.00 % 0.11 % ceph-osd -n osd.24 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] Not a good example, because sometimes mon writes more intensively, but it is very apparent that thread 4919 of the monitor process is the top disk writer in the system. This is the mon thread producing lots of writes: 4919 167 20 0 2031116 1.1g 10652 S 0.0 0.3 288:48.65 rocksdb:low0 Then with a combination of lsof and sysdig I determine that the writes are being made to /var/lib/ceph/mon/ceph-ceph03/store.db/*.sst, i.e. the mon's rocksdb store: ceph-mon 4838 167 200r REG 253,11 67319253 14812899 /var/lib/ceph/mon/ceph-ceph03/store.db/3677146.sst ceph-mon 4838 167 203r REG 253,11 67228736 14813270 /var/lib/ceph/mon/ceph-ceph03/store.db/3677147.sst ceph-mon 4838 167 205r REG 253,11 67243212 14813275 /var/lib/ceph/mon/ceph-ceph03/store.db/3677148.sst ceph-mon 4838 167 208r REG 253,11 67247953 14813316 /var/lib/ceph/mon/ceph-ceph03/store.db/3677149.sst ceph-mon 4838 167 220r REG 253,11 67261659 14813332 /var/lib/ceph/mon/ceph-ceph03/store.db/3677150.sst ceph-mon 4838 167 221r REG 253,11 67242500 14813345 /var/lib/ceph/mon/ceph-ceph03/store.db/3677151.sst ceph-mon 4838 167 224r REG 253,11 67264969 14813348 /var/lib/ceph/mon/ceph-ceph03/store.db/3677152.sst ceph-mon 4838 167 228r REG 253,11 64346933 14813381 /var/lib/ceph/mon/ceph-ceph03/store.db/3677153.sst By matching iotop and sysdig write records to mon's log entries, I see that the writes happen during "manual compaction" events - whatever they are, because there's no documentation on this whatsoever, and each time around 0.56GB is being written to disk to a new set of *.sst files, which is the total size of the store.db. Looks like from time to time the monitor just reads its store.db and writes it out to a new set of files, as the file names "numbers" increase with each write: ceph-mon 4838 167 175r REG 253,11 67220863 14812310 /var/lib/ceph/mon/ceph-ceph03/store.db/3677167.sst ceph-mon 4838 167 200r REG 253,11 67358627 14812899 /var/lib/ceph/mon/ceph-ceph03/store.db/3677168.sst ceph-mon 4838 167 203r REG 253,11 67277978 14813270 /var/lib/ceph/mon/ceph-ceph03/store.db/3677169.sst ceph-mon 4838 167 205r REG 253,11 67256312 14813275 /var/lib/ceph/mon/ceph-ceph03/store.db/3677170.sst ceph-mon 4838 167 208r REG 253,11 67226761 14813316 /var/lib/ceph/mon/ceph-ceph03/store.db/3677171.sst ceph-mon 4838 167 220r REG 253,11 67258798 14813332 /var/lib/ceph/mon/ceph-ceph03/store.db/3677172.sst ceph-mon 4838 167 221r REG 253,11 67224665 14813345 /var/lib/ceph/mon/ceph-ceph03/store.db/3677173.sst ceph-mon 4838 167 224r REG 253,11 67224123 14813348 /var/lib/ceph/mon/ceph-ceph03/store.db/3677174.sst ceph-mon 4838 167 228r REG 253,11 62195349 14813381 /var/lib/ceph/mon/ceph-ceph03/store.db/3677175.sst I hope this clears up the situation. Do you observe this behavior in your clusters? Can you please check whether your mons do something similar and store.db/*.sst change often? /Z On Wed, 11 Oct 2023 at 16:22, Eugen Block <eblock@nde.ag<mailto:eblock@nde.ag>> wrote: That all looks normal to me, to be honest. Can you show some details how you calculate the "hundreds of GB per day"? I see similar stats as Frank on different clusters with different client IO. Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com>>:
Sure, nothing unusual there:
-------
cluster: id: 3f50555a-ae2a-11eb-a2fc-ffde44714d86 health: HEALTH_OK
services: mon: 5 daemons, quorum ceph01,ceph03,ceph04,ceph05,ceph02 (age 2w) mgr: ceph01.vankui(active, since 12d), standbys: ceph02.shsinf osd: 96 osds: 96 up (since 2w), 95 in (since 3w)
data: pools: 10 pools, 2400 pgs objects: 6.23M objects, 16 TiB usage: 61 TiB used, 716 TiB / 777 TiB avail pgs: 2396 active+clean 3 active+clean+scrubbing+deep 1 active+clean+scrubbing
io: client: 2.7 GiB/s rd, 27 MiB/s wr, 46.95k op/s rd, 2.17k op/s wr
-------
Please disregard the big read number, a customer is running a read-intensive job. Mon store writes keep happening when the cluster is much more quiet, thus I think that intensive reads have no effect on the mons.
Mgr:
"always_on_modules": [ "balancer", "crash", "devicehealth", "orchestrator", "pg_autoscaler", "progress", "rbd_support", "status", "telemetry", "volumes" ], "enabled_modules": [ "cephadm", "dashboard", "iostat", "prometheus", "restful" ],
-------
/Z
On Wed, 11 Oct 2023 at 14:50, Eugen Block <eblock@nde.ag<mailto:eblock@nde.ag>> wrote:
Can you add some more details as requested by Frank? Which mgr modules are enabled? What's the current 'ceph -s' output?
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed
Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com>>:
We don't use CephFS at all and don't have RBD snapshots apart from some cloning for Openstack images.
The size of mon stores isn't an issue, it's < 600 MB. But it gets overwritten often causing lots of disk writes, and that is an issue for us.
/Z
On Wed, 11 Oct 2023 at 14:37, Eugen Block <eblock@nde.ag<mailto:eblock@nde.ag>> wrote:
Do you use many snapshots (rbd or cephfs)? That can cause a heavy monitor usage, we've seen large mon stores on customer clusters with rbd mirroring on snapshot basis. In a healthy cluster they have mon stores of around 2GB in size.
@Eugen: Was there not an option to limit logging to the MON store?
I don't recall at the moment, worth checking tough.
Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com>>:
9@6 files to L6 => 594449971 bytes debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760] [default] compacted to: base level 5 level multiplier 10.00 max bytes base 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd, 173.9 wr, level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9), read-write-amplify(3434.6) write-amplify(1707.7) OK, records in: 35108, records dropped: 516 output_compression: NoCompression debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros": 1697022613921002, "job": 70854, "event": "compaction_finished", "compaction_time_micros": 3418822, "compaction_time_cpu_micros": 785454, "output_level": 6, "num_output_files": 9, "total_output_size": 594449971, "num_input_records": 35108, "num_output_records": 34592, "num_subcompactions": 1, "output_compression": "NoCompression", "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, "lsm_state": [0, 0, 0, 0, 0, 0, 9]}
The log even mentions the huge write multiplication. I wonder whether this is normal and what can be done about it.
/Z
On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk<mailto:frans@dtu.dk>> wrote:
I need to ask here: where exactly do you observe the hundreds of GB written per day? Are the mon logs huge? Is it the mon store? Is your cluster unhealthy?
We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about 800 librbd clients. Per week our mon logs are about 70M, the cluster logs about 120M , the audit logs about 70M and I see between 100-200Kb/s writes to the mon store. That's in the lower-digit GB range per day. Hundreds of GB per day sound completely over the top on a healthy cluster, unless you have MGR modules changing the OSD/cluster map continuously.
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed by now?
@Eugen: Was there not an option to limit logging to the MON store?
For information to readers, we followed old recommendations from a Dell white paper for building a ceph cluster and have a 1TB Raid10 array on 6x write intensive SSDs for the MON stores. After 5 years we are below 10% wear. Average size of the MON store for a healthy cluster is 500M-1G, but we have seen this ballooning to 100+GB in degraded conditions.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com>> Sent: Wednesday, October 11, 2023 12:00 PM To: Eugen Block Cc: ceph-users@ceph.io<mailto:ceph-users@ceph.io> Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Thank you, Eugen.
I'm interested specifically to find out whether the huge amount of data written by monitors is expected. It is eating through the endurance of our system drives, which were not specced for high DWPD/TBW, as this is not a documented requirement, and monitors produce hundreds of gigabytes of writes per day. I am looking for ways to reduce the amount of writes, if possible.
/Z
On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag<mailto:eblock@nde.ag>> wrote:
> Hi, > > what you report is the expected behaviour, at least I see the same on > all clusters. I can't answer why the compaction is required that > often, but you can control the log level of the rocksdb output: > > ceph config set mon debug_rocksdb 1/5 (default is 4/5) > > This reduces the log entries and you wouldn't see the manual > compaction logs anymore. There are a couple more rocksdb options but I > probably wouldn't change too much, only if you know what you're doing. > Maybe Igor can comment if some other tuning makes sense here. > > Regards, > Eugen > > Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com>>: > > > Any input from anyone, please? > > > > On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko < zakhar@gmail.com<mailto:zakhar@gmail.com>> > wrote: > > > >> Any input from anyone, please? > >> > >> It's another thing that seems to be rather poorly documented: it's > unclear > >> what to expect, what 'normal' behavior should be, and what can be done > >> about the huge amount of writes by monitors. > >> > >> /Z > >> > >> On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko < zakhar@gmail.com<mailto:zakhar@gmail.com>> > wrote: > >> > >>> Hi, > >>> > >>> Monitors in our 16.2.14 cluster appear to quite often run "manual > >>> compaction" tasks: > >>> > >>> debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: > EVENT_LOG_v1 > >>> {"time_micros": 1696843853892760, "job": 64225, "event": > "flush_started", > >>> "num_memtables": 1, "num_entries": 715, "num_deletes": 251, > >>> "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": > >>> "Manual Compaction"} > >>> debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:30:53.910204) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-0 to level-5 from 'paxos .. > 'paxos; > >>> will stop at (end) > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:30:53.911004) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-5 to level-6 from 'paxos .. > 'paxos; > >>> will stop at (end) > >>> debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: > EVENT_LOG_v1 > >>> {"time_micros": 1696843928961390, "job": 64228, "event": > "flush_started", > >>> "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, > >>> "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": > >>> "Manual Compaction"} > >>> debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:32:08.977739) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-0 to level-5 from 'logm .. > 'logm; > >>> will stop at (end) > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:32:08.978512) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-5 to level-6 from 'logm .. > 'logm; > >>> will stop at (end) > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: > EVENT_LOG_v1 > >>> {"time_micros": 1696844009033151, "job": 64231, "event": > "flush_started", > >>> "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, > >>> "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": > >>> "Manual Compaction"} > >>> debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:33:29.049585) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-0 to level-5 from 'paxos .. > 'paxos; > >>> will stop at (end) > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual compaction > >>> starting > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original > Log > >>> Time 2023/10/09-09:33:29.050355) > [db_impl/db_impl_compaction_flush.cc:2516] > >>> [default] Manual compaction from level-5 to level-6 from 'paxos .. > 'paxos; > >>> will stop at (end) > >>> > >>> I have removed a lot of interim log messages to save space. > >>> > >>> During each compaction the monitor process writes approximately 500-600 > >>> MB of data to disk over a short period of time. These writes add up to > tens > >>> of gigabytes per hour and hundreds of gigabytes per day. > >>> > >>> Monitor rocksdb and compaction options are default: > >>> > >>> "mon_compact_on_bootstrap": "false", > >>> "mon_compact_on_start": "false", > >>> "mon_compact_on_trim": "true", > >>> "mon_rocksdb_options": > >>> >
9@6 files to L6, score -1.00 debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: EVENT_LOG_v1 {"time_micros": 1697022610487624, "job": 70854, "event": "compaction_started", "compaction_reason": "ManualCompaction", "files_L5": [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536, 3675537, 3675538, 3675539, 3675540, 3675541], "score": -1, "input_data_size": 601117031} debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675544: 2015 keys, 67287115 bytes debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675545: 24343 keys, 67336225 bytes debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675546: 1196 keys, 67225813 bytes debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675547: 1049 keys, 67252678 bytes debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675548: 1081 keys, 67216638 bytes debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675549: 1196 keys, 67245376 bytes debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675550: 1195 keys, 67246813 bytes debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675551: 1205 keys, 67223302 bytes debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table #3675552: 1312 keys, 56416011 bytes debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted 1@5
Thank you, Frank.
The cluster is healthy, operating normally, nothing unusual is going on. We observe lots of writes by mon processes into mon rocksdb stores, specifically:
/var/lib/ceph/mon/ceph-cephXX/store.db: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
The site of the files is not huge, but monitors rotate and write out these files often, sometimes several times per minute, resulting in lots of data written to disk. The writes coincide with "manual compaction" events logged by the monitors, for example:
debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: [compaction/compaction_job.cc:1676] [default] [JOB 70854] Compacting 1@5
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
> >>> > >>> Is this expected behavior? Is this something I can adjust in order to > >>> extend the system storage life? > >>> > >>> Best regards, > >>> Zakhar > >>> > >> > > _______________________________________________ > > ceph-users mailing list -- ceph-users@ceph.io<mailto:ceph-users@ceph.io> > > To unsubscribe send an email to ceph-users-leave@ceph.io<mailto:ceph-users-leave@ceph.io> > > > _______________________________________________ > ceph-users mailing list -- ceph-users@ceph.io<mailto:ceph-users@ceph.io> > To unsubscribe send an email to ceph-users-leave@ceph.io<mailto:ceph-users-leave@ceph.io> > _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io<mailto:ceph-users@ceph.io> To unsubscribe send an email to ceph-users-leave@ceph.io<mailto:ceph-users-leave@ceph.io>
Thank you, Frank. This confirms that monitors indeed do this, and
Our boot drives in 3 systems are smaller 1 DWPD drives (RAID1 to protect
against a random single drive failure), and over 3 years mons have eaten
through 60% of their endurance. Other systems have larger boot drives and
2% of their endurance were used up over 1.5 years.
It would still be good to get an understanding why monitors do this, and
whether there is any way to reduce the amount of writes. Unfortunately,
Ceph documentation in this regard is severely lacking.
I'm copying this to ceph-docs, perhaps someone will find it useful and
adjust the hardware recommendations.
/Z
On Wed, 11 Oct 2023, 18:23 Frank Schilder, <frans@dtu.dk> wrote:
> Oh wow! I never bothered looking, because on our hardware the wear is so
> low:
>
> # iotop -ao -bn 2 -d 300
> Total DISK READ : 0.00 B/s | Total DISK WRITE : 6.46 M/s
> Actual DISK READ: 0.00 B/s | Actual DISK WRITE: 6.47 M/s
> TID PRIO USER DISK READ DISK WRITE SWAPIN IO COMMAND
> 2230 be/4 ceph 0.00 B 1818.71 M 0.00 % 0.46 % ceph-mon
> --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01
> --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65
> [rocksdb:low0]
> 2256 be/4 ceph 0.00 B 19.27 M 0.00 % 0.43 % ceph-mon
> --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01
> --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65
> [safe_timer]
> 2250 be/4 ceph 0.00 B 42.38 M 0.00 % 0.26 % ceph-mon
> --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01
> --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65
> [fn_monstore]
> 2231 be/4 ceph 0.00 B 58.36 M 0.00 % 0.01 % ceph-mon
> --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01
> --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65
> [rocksdb:high0]
> 644 be/3 root 0.00 B 576.00 K 0.00 % 0.00 % [jbd2/sda3-8]
> 2225 be/4 ceph 0.00 B 128.00 K 0.00 % 0.00 % ceph-mon
> --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01
> --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [log]
> 1637141 be/4 root 0.00 B 0.00 B 0.00 % 0.00 %
> [kworker/u113:2-flush-8:0]
> 1636453 be/4 root 0.00 B 0.00 B 0.00 % 0.00 %
> [kworker/u112:0-ceph0]
> 1560 be/4 root 0.00 B 20.00 K 0.00 % 0.00 % rsyslogd -n
> [in:imjournal]
> 1561 be/4 root 0.00 B 56.00 K 0.00 % 0.00 % rsyslogd -n
> [rs:main Q:Reg]
>
> 1.8GB every 5 minutes, thats 518GB per day. The 400G drives we have are
> rated 10DWPD and with the 6-drives RAID10 config this gives plenty of
> life-time. I guess this write load will kill any low-grade SSD (typical
> bood devices, even enterprise ones) specifically if its smaller drives and
> the controller doesn't reallocate cells according to remaining write
> endurance.
>
> I guess there was a reason for the recommendations by Dell. I always
> thought that the recent recommendation for MON store storage in the ceph
> docs are a "bit unrealistic", apparently both, in size and in performance
> (including endurance). Well, I guess you need to look for write intensive
> drives with decent specs. If you do, also go for sufficient size. This will
> absorb temporary usage peaks that can be very large and also provide extra
> endurance with SSDs with good controllers.
>
> I also think the recommendations on the ceph docs deserve a reality check.
>
> Best regards,
> =================
> Frank Schilder
> AIT Risø Campus
> Bygning 109, rum S14
>
> ________________________________________
> From: Zakhar Kirpichenko <zakhar@gmail.com>
> Sent: Wednesday, October 11, 2023 4:30 PM
> To: Eugen Block
> Cc: Frank Schilder; ceph-users@ceph.io
> Subject: Re: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
>
> Eugen,
>
> Thanks for your response. May I ask what numbers you're referring to?
>
> I am not referring to monitor store.db sizes. I am specifically referring
> to writes monitors do to their store.db file by frequently rotating and
> replacing them with new versions during compactions. The size of the
> store.db remains more or less the same.
>
> This is a 300s iotop snippet, sorted by aggregated disk writes:
>
> Total DISK READ: 35.56 M/s | Total DISK WRITE: 23.89 M/s
> Current DISK READ: 35.64 M/s | Current DISK WRITE: 24.09 M/s
> TID PRIO USER DISK READ DISK WRITE> SWAPIN IO COMMAND
> 4919 be/4 167 16.75 M 2.24 G 0.00 % 1.34 % ceph-mon -n
> mon.ceph03 -f --setuser ceph --setgr~lt-mon-cluster-log-to-stderr=true
> [rocksdb:low0]
> 15122 be/4 167 0.00 B 652.91 M 0.00 % 0.27 % ceph-osd -n
> osd.31 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 17073 be/4 167 0.00 B 651.86 M 0.00 % 0.27 % ceph-osd -n
> osd.32 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 17268 be/4 167 0.00 B 490.86 M 0.00 % 0.18 % ceph-osd -n
> osd.25 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 18032 be/4 167 0.00 B 463.57 M 0.00 % 0.17 % ceph-osd -n
> osd.26 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 16855 be/4 167 0.00 B 402.86 M 0.00 % 0.15 % ceph-osd -n
> osd.22 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 17406 be/4 167 0.00 B 387.03 M 0.00 % 0.14 % ceph-osd -n
> osd.27 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 17932 be/4 167 0.00 B 375.42 M 0.00 % 0.13 % ceph-osd -n
> osd.29 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 18017 be/4 167 0.00 B 359.38 M 0.00 % 0.13 % ceph-osd -n
> osd.28 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 17420 be/4 167 0.00 B 332.83 M 0.00 % 0.12 % ceph-osd -n
> osd.23 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 17975 be/4 167 0.00 B 312.06 M 0.00 % 0.11 % ceph-osd -n
> osd.30 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
> 17273 be/4 167 0.00 B 303.49 M 0.00 % 0.11 % ceph-osd -n
> osd.24 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug
> [bstore_kv_sync]
>
> Not a good example, because sometimes mon writes more intensively, but it
> is very apparent that thread 4919 of the monitor process is the top disk
> writer in the system.
>
> This is the mon thread producing lots of writes:
>
> 4919 167 20 0 2031116 1.1g 10652 S 0.0 0.3 288:48.65
> rocksdb:low0
>
> Then with a combination of lsof and sysdig I determine that the writes are
> being made to /var/lib/ceph/mon/ceph-ceph03/store.db/*.sst, i.e. the mon's
> rocksdb store:
>
> ceph-mon 4838 167 200r REG 253,11 67319253 14812899
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677146.sst
> ceph-mon 4838 167 203r REG 253,11 67228736 14813270
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677147.sst
> ceph-mon 4838 167 205r REG 253,11 67243212 14813275
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677148.sst
> ceph-mon 4838 167 208r REG 253,11 67247953 14813316
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677149.sst
> ceph-mon 4838 167 220r REG 253,11 67261659 14813332
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677150.sst
> ceph-mon 4838 167 221r REG 253,11 67242500 14813345
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677151.sst
> ceph-mon 4838 167 224r REG 253,11 67264969 14813348
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677152.sst
> ceph-mon 4838 167 228r REG 253,11 64346933 14813381
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677153.sst
>
> By matching iotop and sysdig write records to mon's log entries, I see
> that the writes happen during "manual compaction" events - whatever they
> are, because there's no documentation on this whatsoever, and each time
> around 0.56GB is being written to disk to a new set of *.sst files, which
> is the total size of the store.db. Looks like from time to time the monitor
> just reads its store.db and writes it out to a new set of files, as the
> file names "numbers" increase with each write:
>
> ceph-mon 4838 167 175r REG 253,11 67220863 14812310
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677167.sst
> ceph-mon 4838 167 200r REG 253,11 67358627 14812899
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677168.sst
> ceph-mon 4838 167 203r REG 253,11 67277978 14813270
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677169.sst
> ceph-mon 4838 167 205r REG 253,11 67256312 14813275
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677170.sst
> ceph-mon 4838 167 208r REG 253,11 67226761 14813316
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677171.sst
> ceph-mon 4838 167 220r REG 253,11 67258798 14813332
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677172.sst
> ceph-mon 4838 167 221r REG 253,11 67224665 14813345
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677173.sst
> ceph-mon 4838 167 224r REG 253,11 67224123 14813348
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677174.sst
> ceph-mon 4838 167 228r REG 253,11 62195349 14813381
> /var/lib/ceph/mon/ceph-ceph03/store.db/3677175.sst
>
> I hope this clears up the situation.
>
> Do you observe this behavior in your clusters? Can you please check
> whether your mons do something similar and store.db/*.sst change often?
>
> /Z
>
> On Wed, 11 Oct 2023 at 16:22, Eugen Block <eblock@nde.ag<mailto:
> eblock@nde.ag>> wrote:
> That all looks normal to me, to be honest. Can you show some details
> how you calculate the "hundreds of GB per day"? I see similar stats as
> Frank on different clusters with different client IO.
>
> Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com>>:
>
> > Sure, nothing unusual there:
> >
> > -------
> >
> > cluster:
> > id: 3f50555a-ae2a-11eb-a2fc-ffde44714d86
> > health: HEALTH_OK
> >
> > services:
> > mon: 5 daemons, quorum ceph01,ceph03,ceph04,ceph05,ceph02 (age 2w)
> > mgr: ceph01.vankui(active, since 12d), standbys: ceph02.shsinf
> > osd: 96 osds: 96 up (since 2w), 95 in (since 3w)
> >
> > data:
> > pools: 10 pools, 2400 pgs
> > objects: 6.23M objects, 16 TiB
> > usage: 61 TiB used, 716 TiB / 777 TiB avail
> > pgs: 2396 active+clean
> > 3 active+clean+scrubbing+deep
> > 1 active+clean+scrubbing
> >
> > io:
> > client: 2.7 GiB/s rd, 27 MiB/s wr, 46.95k op/s rd, 2.17k op/s wr
> >
> > -------
> >
> > Please disregard the big read number, a customer is running a
> > read-intensive job. Mon store writes keep happening when the cluster is
> > much more quiet, thus I think that intensive reads have no effect on the
> > mons.
> >
> > Mgr:
> >
> > "always_on_modules": [
> > "balancer",
> > "crash",
> > "devicehealth",
> > "orchestrator",
> > "pg_autoscaler",
> > "progress",
> > "rbd_support",
> > "status",
> > "telemetry",
> > "volumes"
> > ],
> > "enabled_modules": [
> > "cephadm",
> > "dashboard",
> > "iostat",
> > "prometheus",
> > "restful"
> > ],
> >
> > -------
> >
> > /Z
> >
> >
> > On Wed, 11 Oct 2023 at 14:50, Eugen Block <eblock@nde.ag<mailto:
> eblock@nde.ag>> wrote:
> >
> >> Can you add some more details as requested by Frank? Which mgr modules
> >> are enabled? What's the current 'ceph -s' output?
> >>
> >> > Is autoscaler running and doing stuff?
> >> > Is balancer running and doing stuff?
> >> > Is backfill going on?
> >> > Is recovery going on?
> >> > Is your ceph version affected by the "excessive logging to MON
> >> > store" issue that was present starting with pacific but should have
> >> > been addressed
> >>
> >>
> >> Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com
> >>:
> >>
> >> > We don't use CephFS at all and don't have RBD snapshots apart from
> some
> >> > cloning for Openstack images.
> >> >
> >> > The size of mon stores isn't an issue, it's < 600 MB. But it gets
> >> > overwritten often causing lots of disk writes, and that is an issue
> for
> >> us.
> >> >
> >> > /Z
> >> >
> >> > On Wed, 11 Oct 2023 at 14:37, Eugen Block <eblock@nde.ag<mailto:
> eblock@nde.ag>> wrote:
> >> >
> >> >> Do you use many snapshots (rbd or cephfs)? That can cause a heavy
> >> >> monitor usage, we've seen large mon stores on customer clusters with
> >> >> rbd mirroring on snapshot basis. In a healthy cluster they have mon
> >> >> stores of around 2GB in size.
> >> >>
> >> >> >> @Eugen: Was there not an option to limit logging to the MON store?
> >> >>
> >> >> I don't recall at the moment, worth checking tough.
> >> >>
> >> >> Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:
> zakhar@gmail.com>>:
> >> >>
> >> >> > Thank you, Frank.
> >> >> >
> >> >> > The cluster is healthy, operating normally, nothing unusual is
> going
> >> on.
> >> >> We
> >> >> > observe lots of writes by mon processes into mon rocksdb stores,
> >> >> > specifically:
> >> >> >
> >> >> > /var/lib/ceph/mon/ceph-cephXX/store.db:
> >> >> > 65M 3675511.sst
> >> >> > 65M 3675512.sst
> >> >> > 65M 3675513.sst
> >> >> > 65M 3675514.sst
> >> >> > 65M 3675515.sst
> >> >> > 65M 3675516.sst
> >> >> > 65M 3675517.sst
> >> >> > 65M 3675518.sst
> >> >> > 62M 3675519.sst
> >> >> >
> >> >> > The site of the files is not huge, but monitors rotate and write
> out
> >> >> these
> >> >> > files often, sometimes several times per minute, resulting in lots
> of
> >> >> data
> >> >> > written to disk. The writes coincide with "manual compaction"
> events
> >> >> logged
> >> >> > by the monitors, for example:
> >> >> >
> >> >> > debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1676] [default] [JOB 70854]
> Compacting
> >> 1@5
> >> >> +
> >> >> > 9@6 files to L6, score -1.00
> >> >> > debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb:
> >> EVENT_LOG_v1
> >> >> > {"time_micros": 1697022610487624, "job": 70854, "event":
> >> >> > "compaction_started", "compaction_reason": "ManualCompaction",
> >> >> "files_L5":
> >> >> > [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536,
> 3675537,
> >> >> > 3675538, 3675539, 3675540, 3675541], "score": -1,
> "input_data_size":
> >> >> > 601117031}
> >> >> > debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675544: 2015 keys, 67287115 bytes
> >> >> > debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675545: 24343 keys, 67336225 bytes
> >> >> > debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675546: 1196 keys, 67225813 bytes
> >> >> > debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675547: 1049 keys, 67252678 bytes
> >> >> > debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675548: 1081 keys, 67216638 bytes
> >> >> > debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675549: 1196 keys, 67245376 bytes
> >> >> > debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675550: 1195 keys, 67246813 bytes
> >> >> > debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675551: 1205 keys, 67223302 bytes
> >> >> > debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated
> >> table
> >> >> > #3675552: 1312 keys, 56416011 bytes
> >> >> > debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb:
> >> >> > [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted
> >> 1@5
> >> >> +
> >> >> > 9@6 files to L6 => 594449971 bytes
> >> >> > debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb:
> (Original
> >> Log
> >> >> > Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760]
> >> >> > [default] compacted to: base level 5 level multiplier 10.00 max
> bytes
> >> >> base
> >> >> > 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd,
> 173.9
> >> >> wr,
> >> >> > level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9),
> >> >> > read-write-amplify(3434.6) write-amplify(1707.7) OK, records in:
> >> 35108,
> >> >> > records dropped: 516 output_compression: NoCompression
> >> >> > debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb:
> (Original
> >> Log
> >> >> > Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros":
> >> >> > 1697022613921002, "job": 70854, "event": "compaction_finished",
> >> >> > "compaction_time_micros": 3418822, "compaction_time_cpu_micros":
> >> 785454,
> >> >> > "output_level": 6, "num_output_files": 9, "total_output_size":
> >> 594449971,
> >> >> > "num_input_records": 35108, "num_output_records": 34592,
> >> >> > "num_subcompactions": 1, "output_compression": "NoCompression",
> >> >> > "num_single_delete_mismatches": 0,
> "num_single_delete_fallthrough": 0,
> >> >> > "lsm_state": [0, 0, 0, 0, 0, 0, 9]}
> >> >> >
> >> >> > The log even mentions the huge write multiplication. I wonder
> whether
> >> >> this
> >> >> > is normal and what can be done about it.
> >> >> >
> >> >> > /Z
> >> >> >
> >> >> > On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk<mailto:
> frans@dtu.dk>> wrote:
> >> >> >
> >> >> >> I need to ask here: where exactly do you observe the hundreds of
> GB
> >> >> >> written per day? Are the mon logs huge? Is it the mon store? Is
> your
> >> >> >> cluster unhealthy?
> >> >> >>
> >> >> >> We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients
> and
> >> >> about
> >> >> >> 800 librbd clients. Per week our mon logs are about 70M, the
> cluster
> >> >> logs
> >> >> >> about 120M , the audit logs about 70M and I see between
> 100-200Kb/s
> >> >> writes
> >> >> >> to the mon store. That's in the lower-digit GB range per day.
> >> Hundreds
> >> >> of
> >> >> >> GB per day sound completely over the top on a healthy cluster,
> unless
> >> >> you
> >> >> >> have MGR modules changing the OSD/cluster map continuously.
> >> >> >>
> >> >> >> Is autoscaler running and doing stuff?
> >> >> >> Is balancer running and doing stuff?
> >> >> >> Is backfill going on?
> >> >> >> Is recovery going on?
> >> >> >> Is your ceph version affected by the "excessive logging to MON
> store"
> >> >> >> issue that was present starting with pacific but should have been
> >> >> addressed
> >> >> >> by now?
> >> >> >>
> >> >> >> @Eugen: Was there not an option to limit logging to the MON store?
> >> >> >>
> >> >> >> For information to readers, we followed old recommendations from a
> >> Dell
> >> >> >> white paper for building a ceph cluster and have a 1TB Raid10
> array
> >> on
> >> >> 6x
> >> >> >> write intensive SSDs for the MON stores. After 5 years we are
> below
> >> 10%
> >> >> >> wear. Average size of the MON store for a healthy cluster is
> 500M-1G,
> >> >> but
> >> >> >> we have seen this ballooning to 100+GB in degraded conditions.
> >> >> >>
> >> >> >> Best regards,
> >> >> >> =================
> >> >> >> Frank Schilder
> >> >> >> AIT Risø Campus
> >> >> >> Bygning 109, rum S14
> >> >> >>
> >> >> >> ________________________________________
> >> >> >> From: Zakhar Kirpichenko <zakhar@gmail.com<mailto:
> zakhar@gmail.com>>
> >> >> >> Sent: Wednesday, October 11, 2023 12:00 PM
> >> >> >> To: Eugen Block
> >> >> >> Cc: ceph-users@ceph.io<mailto:ceph-users@ceph.io>
> >> >> >> Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
> >> >> >>
> >> >> >> Thank you, Eugen.
> >> >> >>
> >> >> >> I'm interested specifically to find out whether the huge amount of
> >> data
> >> >> >> written by monitors is expected. It is eating through the
> endurance
> >> of
> >> >> our
> >> >> >> system drives, which were not specced for high DWPD/TBW, as this
> is
> >> not
> >> >> a
> >> >> >> documented requirement, and monitors produce hundreds of
> gigabytes of
> >> >> >> writes per day. I am looking for ways to reduce the amount of
> >> writes, if
> >> >> >> possible.
> >> >> >>
> >> >> >> /Z
> >> >> >>
> >> >> >> On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag<mailto:
> eblock@nde.ag>> wrote:
> >> >> >>
> >> >> >> > Hi,
> >> >> >> >
> >> >> >> > what you report is the expected behaviour, at least I see the
> same
> >> on
> >> >> >> > all clusters. I can't answer why the compaction is required that
> >> >> >> > often, but you can control the log level of the rocksdb output:
> >> >> >> >
> >> >> >> > ceph config set mon debug_rocksdb 1/5 (default is 4/5)
> >> >> >> >
> >> >> >> > This reduces the log entries and you wouldn't see the manual
> >> >> >> > compaction logs anymore. There are a couple more rocksdb options
> >> but I
> >> >> >> > probably wouldn't change too much, only if you know what you're
> >> doing.
> >> >> >> > Maybe Igor can comment if some other tuning makes sense here.
> >> >> >> >
> >> >> >> > Regards,
> >> >> >> > Eugen
> >> >> >> >
> >> >> >> > Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:
> zakhar@gmail.com>>:
> >> >> >> >
> >> >> >> > > Any input from anyone, please?
> >> >> >> > >
> >> >> >> > > On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko <
> >> zakhar@gmail.com<mailto:zakhar@gmail.com>>
> >> >> >> > wrote:
> >> >> >> > >
> >> >> >> > >> Any input from anyone, please?
> >> >> >> > >>
> >> >> >> > >> It's another thing that seems to be rather poorly documented:
> >> it's
> >> >> >> > unclear
> >> >> >> > >> what to expect, what 'normal' behavior should be, and what
> can
> >> be
> >> >> done
> >> >> >> > >> about the huge amount of writes by monitors.
> >> >> >> > >>
> >> >> >> > >> /Z
> >> >> >> > >>
> >> >> >> > >> On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko <
> >> zakhar@gmail.com<mailto:zakhar@gmail.com>>
> >> >> >> > wrote:
> >> >> >> > >>
> >> >> >> > >>> Hi,
> >> >> >> > >>>
> >> >> >> > >>> Monitors in our 16.2.14 cluster appear to quite often run
> >> "manual
> >> >> >> > >>> compaction" tasks:
> >> >> >> > >>>
> >> >> >> > >>> debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb:
> >> >> >> > EVENT_LOG_v1
> >> >> >> > >>> {"time_micros": 1696843853892760, "job": 64225, "event":
> >> >> >> > "flush_started",
> >> >> >> > >>> "num_memtables": 1, "num_entries": 715, "num_deletes": 251,
> >> >> >> > >>> "total_data_size": 3870352, "memory_usage": 3886744,
> >> >> "flush_reason":
> >> >> >> > >>> "Manual Compaction"}
> >> >> >> > >>> debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:30:53.910204)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-0 to level-5 from
> >> 'paxos ..
> >> >> >> > 'paxos;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:30:53.911004)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-5 to level-6 from
> >> 'paxos ..
> >> >> >> > 'paxos;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb:
> >> >> >> > EVENT_LOG_v1
> >> >> >> > >>> {"time_micros": 1696843928961390, "job": 64228, "event":
> >> >> >> > "flush_started",
> >> >> >> > >>> "num_memtables": 1, "num_entries": 1580, "num_deletes": 502,
> >> >> >> > >>> "total_data_size": 8404605, "memory_usage": 8465840,
> >> >> "flush_reason":
> >> >> >> > >>> "Manual Compaction"}
> >> >> >> > >>> debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:32:08.977739)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-0 to level-5 from
> 'logm
> >> ..
> >> >> >> > 'logm;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:32:08.978512)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-5 to level-6 from
> 'logm
> >> ..
> >> >> >> > 'logm;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb:
> >> >> >> > EVENT_LOG_v1
> >> >> >> > >>> {"time_micros": 1696844009033151, "job": 64231, "event":
> >> >> >> > "flush_started",
> >> >> >> > >>> "num_memtables": 1, "num_entries": 1430, "num_deletes": 251,
> >> >> >> > >>> "total_data_size": 8975535, "memory_usage": 9035920,
> >> >> "flush_reason":
> >> >> >> > >>> "Manual Compaction"}
> >> >> >> > >>> debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:33:29.049585)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-0 to level-5 from
> >> 'paxos ..
> >> >> >> > 'paxos;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb:
> >> >> >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual
> >> >> >> compaction
> >> >> >> > >>> starting
> >> >> >> > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb:
> >> >> (Original
> >> >> >> > Log
> >> >> >> > >>> Time 2023/10/09-09:33:29.050355)
> >> >> >> > [db_impl/db_impl_compaction_flush.cc:2516]
> >> >> >> > >>> [default] Manual compaction from level-5 to level-6 from
> >> 'paxos ..
> >> >> >> > 'paxos;
> >> >> >> > >>> will stop at (end)
> >> >> >> > >>>
> >> >> >> > >>> I have removed a lot of interim log messages to save space.
> >> >> >> > >>>
> >> >> >> > >>> During each compaction the monitor process writes
> approximately
> >> >> >> 500-600
> >> >> >> > >>> MB of data to disk over a short period of time. These writes
> >> add
> >> >> up
> >> >> >> to
> >> >> >> > tens
> >> >> >> > >>> of gigabytes per hour and hundreds of gigabytes per day.
> >> >> >> > >>>
> >> >> >> > >>> Monitor rocksdb and compaction options are default:
> >> >> >> > >>>
> >> >> >> > >>> "mon_compact_on_bootstrap": "false",
> >> >> >> > >>> "mon_compact_on_start": "false",
> >> >> >> > >>> "mon_compact_on_trim": "true",
> >> >> >> > >>> "mon_rocksdb_options":
> >> >> >> > >>>
> >> >> >> >
> >> >> >>
> >> >>
> >>
> "write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
> >> >> >> > >>>
> >> >> >> > >>> Is this expected behavior? Is this something I can adjust in
> >> >> order to
> >> >> >> > >>> extend the system storage life?
> >> >> >> > >>>
> >> >> >> > >>> Best regards,
> >> >> >> > >>> Zakhar
> >> >> >> > >>>
> >> >> >> > >>
> >> >> >> > > _______________________________________________
> >> >> >> > > ceph-users mailing list -- ceph-users@ceph.io<mailto:
> ceph-users@ceph.io>
> >> >> >> > > To unsubscribe send an email to ceph-users-leave@ceph.io
> <mailto:ceph-users-leave@ceph.io>
> >> >> >> >
> >> >> >> >
> >> >> >> > _______________________________________________
> >> >> >> > ceph-users mailing list -- ceph-users@ceph.io<mailto:
> ceph-users@ceph.io>
> >> >> >> > To unsubscribe send an email to ceph-users-leave@ceph.io
> <mailto:ceph-users-leave@ceph.io>
> >> >> >> >
> >> >> >> _______________________________________________
> >> >> >> ceph-users mailing list -- ceph-users@ceph.io<mailto:
> ceph-users@ceph.io>
> >> >> >> To unsubscribe send an email to ceph-users-leave@ceph.io<mailto:
> ceph-users-leave@ceph.io>
> >> >> >>
> >> >>
> >> >>
> >> >>
> >> >>
> >>
> >>
> >>
> >>
>
>
>
>
An interesting find: I reduced the number of PGs for some of the less utilized pools, which brought the total number of PGs in the cluster down from 2400 to 1664. The cluster is healthy, the only change was the 30% reduction of PGs, but mons now have a much smaller store.db, have much fewer "manual compaction" events and write significantly less data. Store.db is about 1/2 smaller, ~260 MB in 3 .sst files compared to 590-600 MB in 9 .sst files before the PG changes: total 258012 drwxr-xr-x 2 167 167 4096 Oct 13 16:15 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst -rw-r--r-- 1 167 167 17 Sep 18 11:53 CURRENT -rw-r--r-- 1 167 167 37 Nov 3 2021 IDENTITY -rw-r--r-- 1 167 167 0 Nov 3 2021 LOCK -rw-r--r-- 1 167 167 27039408 Oct 13 16:15 MANIFEST-2785821 -rw-r--r-- 1 167 167 5287 Sep 1 04:39 OPTIONS-2710412 -rw-r--r-- 1 167 167 5287 Sep 18 11:53 OPTIONS-2785824 "Manual compaction" events now run half as often compared to before the change. Before the change, compaction events per hour: # docker logs a4615a23b4c6 2>&1| grep -i 2023-10-13T10 | grep -ci "manual compaction from" 88 After the change, compaction events per hour: # docker logs a4615a23b4c6 2>&1| grep -i 2023-10-13T15 | grep -ci "manual compaction from" 45 I ran several iotop measurements, mons consistently write 550-750 MB to disk every 5 minutes compared to 1.5-2.5 GB every 5 min before the changes: 4919 be/4 167 7.29 M 754.04 M 0.00 % 0.17 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 4919 be/4 167 8.12 M 554.53 M 0.00 % 0.12 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 4919 be/4 167 532.00 K 750.40 M 0.00 % 0.16 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] It is a significant reduction of store.db and associated disk writes. I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in. /Z On Wed, 11 Oct 2023 at 19:00, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Frank. This confirms that monitors indeed do this, and
Our boot drives in 3 systems are smaller 1 DWPD drives (RAID1 to protect against a random single drive failure), and over 3 years mons have eaten through 60% of their endurance. Other systems have larger boot drives and 2% of their endurance were used up over 1.5 years.
It would still be good to get an understanding why monitors do this, and whether there is any way to reduce the amount of writes. Unfortunately, Ceph documentation in this regard is severely lacking.
I'm copying this to ceph-docs, perhaps someone will find it useful and adjust the hardware recommendations.
/Z
On Wed, 11 Oct 2023, 18:23 Frank Schilder, <frans@dtu.dk> wrote:
Oh wow! I never bothered looking, because on our hardware the wear is so low:
# iotop -ao -bn 2 -d 300 Total DISK READ : 0.00 B/s | Total DISK WRITE : 6.46 M/s Actual DISK READ: 0.00 B/s | Actual DISK WRITE: 6.47 M/s TID PRIO USER DISK READ DISK WRITE SWAPIN IO COMMAND 2230 be/4 ceph 0.00 B 1818.71 M 0.00 % 0.46 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [rocksdb:low0] 2256 be/4 ceph 0.00 B 19.27 M 0.00 % 0.43 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [safe_timer] 2250 be/4 ceph 0.00 B 42.38 M 0.00 % 0.26 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [fn_monstore] 2231 be/4 ceph 0.00 B 58.36 M 0.00 % 0.01 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [rocksdb:high0] 644 be/3 root 0.00 B 576.00 K 0.00 % 0.00 % [jbd2/sda3-8] 2225 be/4 ceph 0.00 B 128.00 K 0.00 % 0.00 % ceph-mon --cluster ceph --setuser ceph --setgroup ceph --foreground -i ceph-01 --mon-data /var/lib/ceph/mon/ceph-ceph-01 --public-addr 192.168.32.65 [log] 1637141 be/4 root 0.00 B 0.00 B 0.00 % 0.00 % [kworker/u113:2-flush-8:0] 1636453 be/4 root 0.00 B 0.00 B 0.00 % 0.00 % [kworker/u112:0-ceph0] 1560 be/4 root 0.00 B 20.00 K 0.00 % 0.00 % rsyslogd -n [in:imjournal] 1561 be/4 root 0.00 B 56.00 K 0.00 % 0.00 % rsyslogd -n [rs:main Q:Reg]
1.8GB every 5 minutes, thats 518GB per day. The 400G drives we have are rated 10DWPD and with the 6-drives RAID10 config this gives plenty of life-time. I guess this write load will kill any low-grade SSD (typical bood devices, even enterprise ones) specifically if its smaller drives and the controller doesn't reallocate cells according to remaining write endurance.
I guess there was a reason for the recommendations by Dell. I always thought that the recent recommendation for MON store storage in the ceph docs are a "bit unrealistic", apparently both, in size and in performance (including endurance). Well, I guess you need to look for write intensive drives with decent specs. If you do, also go for sufficient size. This will absorb temporary usage peaks that can be very large and also provide extra endurance with SSDs with good controllers.
I also think the recommendations on the ceph docs deserve a reality check.
Best regards, ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 11, 2023 4:30 PM To: Eugen Block Cc: Frank Schilder; ceph-users@ceph.io Subject: Re: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Eugen,
Thanks for your response. May I ask what numbers you're referring to?
I am not referring to monitor store.db sizes. I am specifically referring to writes monitors do to their store.db file by frequently rotating and replacing them with new versions during compactions. The size of the store.db remains more or less the same.
This is a 300s iotop snippet, sorted by aggregated disk writes:
Total DISK READ: 35.56 M/s | Total DISK WRITE: 23.89 M/s Current DISK READ: 35.64 M/s | Current DISK WRITE: 24.09 M/s TID PRIO USER DISK READ DISK WRITE> SWAPIN IO COMMAND 4919 be/4 167 16.75 M 2.24 G 0.00 % 1.34 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgr~lt-mon-cluster-log-to-stderr=true [rocksdb:low0] 15122 be/4 167 0.00 B 652.91 M 0.00 % 0.27 % ceph-osd -n osd.31 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17073 be/4 167 0.00 B 651.86 M 0.00 % 0.27 % ceph-osd -n osd.32 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17268 be/4 167 0.00 B 490.86 M 0.00 % 0.18 % ceph-osd -n osd.25 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 18032 be/4 167 0.00 B 463.57 M 0.00 % 0.17 % ceph-osd -n osd.26 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 16855 be/4 167 0.00 B 402.86 M 0.00 % 0.15 % ceph-osd -n osd.22 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17406 be/4 167 0.00 B 387.03 M 0.00 % 0.14 % ceph-osd -n osd.27 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17932 be/4 167 0.00 B 375.42 M 0.00 % 0.13 % ceph-osd -n osd.29 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 18017 be/4 167 0.00 B 359.38 M 0.00 % 0.13 % ceph-osd -n osd.28 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17420 be/4 167 0.00 B 332.83 M 0.00 % 0.12 % ceph-osd -n osd.23 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17975 be/4 167 0.00 B 312.06 M 0.00 % 0.11 % ceph-osd -n osd.30 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync] 17273 be/4 167 0.00 B 303.49 M 0.00 % 0.11 % ceph-osd -n osd.24 -f --setuser ceph --setgroup ~default-log-stderr-prefix=debug [bstore_kv_sync]
Not a good example, because sometimes mon writes more intensively, but it is very apparent that thread 4919 of the monitor process is the top disk writer in the system.
This is the mon thread producing lots of writes:
4919 167 20 0 2031116 1.1g 10652 S 0.0 0.3 288:48.65 rocksdb:low0
Then with a combination of lsof and sysdig I determine that the writes are being made to /var/lib/ceph/mon/ceph-ceph03/store.db/*.sst, i.e. the mon's rocksdb store:
ceph-mon 4838 167 200r REG 253,11 67319253 14812899 /var/lib/ceph/mon/ceph-ceph03/store.db/3677146.sst ceph-mon 4838 167 203r REG 253,11 67228736 14813270 /var/lib/ceph/mon/ceph-ceph03/store.db/3677147.sst ceph-mon 4838 167 205r REG 253,11 67243212 14813275 /var/lib/ceph/mon/ceph-ceph03/store.db/3677148.sst ceph-mon 4838 167 208r REG 253,11 67247953 14813316 /var/lib/ceph/mon/ceph-ceph03/store.db/3677149.sst ceph-mon 4838 167 220r REG 253,11 67261659 14813332 /var/lib/ceph/mon/ceph-ceph03/store.db/3677150.sst ceph-mon 4838 167 221r REG 253,11 67242500 14813345 /var/lib/ceph/mon/ceph-ceph03/store.db/3677151.sst ceph-mon 4838 167 224r REG 253,11 67264969 14813348 /var/lib/ceph/mon/ceph-ceph03/store.db/3677152.sst ceph-mon 4838 167 228r REG 253,11 64346933 14813381 /var/lib/ceph/mon/ceph-ceph03/store.db/3677153.sst
By matching iotop and sysdig write records to mon's log entries, I see that the writes happen during "manual compaction" events - whatever they are, because there's no documentation on this whatsoever, and each time around 0.56GB is being written to disk to a new set of *.sst files, which is the total size of the store.db. Looks like from time to time the monitor just reads its store.db and writes it out to a new set of files, as the file names "numbers" increase with each write:
ceph-mon 4838 167 175r REG 253,11 67220863 14812310 /var/lib/ceph/mon/ceph-ceph03/store.db/3677167.sst ceph-mon 4838 167 200r REG 253,11 67358627 14812899 /var/lib/ceph/mon/ceph-ceph03/store.db/3677168.sst ceph-mon 4838 167 203r REG 253,11 67277978 14813270 /var/lib/ceph/mon/ceph-ceph03/store.db/3677169.sst ceph-mon 4838 167 205r REG 253,11 67256312 14813275 /var/lib/ceph/mon/ceph-ceph03/store.db/3677170.sst ceph-mon 4838 167 208r REG 253,11 67226761 14813316 /var/lib/ceph/mon/ceph-ceph03/store.db/3677171.sst ceph-mon 4838 167 220r REG 253,11 67258798 14813332 /var/lib/ceph/mon/ceph-ceph03/store.db/3677172.sst ceph-mon 4838 167 221r REG 253,11 67224665 14813345 /var/lib/ceph/mon/ceph-ceph03/store.db/3677173.sst ceph-mon 4838 167 224r REG 253,11 67224123 14813348 /var/lib/ceph/mon/ceph-ceph03/store.db/3677174.sst ceph-mon 4838 167 228r REG 253,11 62195349 14813381 /var/lib/ceph/mon/ceph-ceph03/store.db/3677175.sst
I hope this clears up the situation.
Do you observe this behavior in your clusters? Can you please check whether your mons do something similar and store.db/*.sst change often?
/Z
On Wed, 11 Oct 2023 at 16:22, Eugen Block <eblock@nde.ag<mailto: eblock@nde.ag>> wrote: That all looks normal to me, to be honest. Can you show some details how you calculate the "hundreds of GB per day"? I see similar stats as Frank on different clusters with different client IO.
Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com>>:
Sure, nothing unusual there:
-------
cluster: id: 3f50555a-ae2a-11eb-a2fc-ffde44714d86 health: HEALTH_OK
services: mon: 5 daemons, quorum ceph01,ceph03,ceph04,ceph05,ceph02 (age 2w) mgr: ceph01.vankui(active, since 12d), standbys: ceph02.shsinf osd: 96 osds: 96 up (since 2w), 95 in (since 3w)
data: pools: 10 pools, 2400 pgs objects: 6.23M objects, 16 TiB usage: 61 TiB used, 716 TiB / 777 TiB avail pgs: 2396 active+clean 3 active+clean+scrubbing+deep 1 active+clean+scrubbing
io: client: 2.7 GiB/s rd, 27 MiB/s wr, 46.95k op/s rd, 2.17k op/s wr
-------
Please disregard the big read number, a customer is running a read-intensive job. Mon store writes keep happening when the cluster is much more quiet, thus I think that intensive reads have no effect on the mons.
Mgr:
"always_on_modules": [ "balancer", "crash", "devicehealth", "orchestrator", "pg_autoscaler", "progress", "rbd_support", "status", "telemetry", "volumes" ], "enabled_modules": [ "cephadm", "dashboard", "iostat", "prometheus", "restful" ],
-------
/Z
On Wed, 11 Oct 2023 at 14:50, Eugen Block <eblock@nde.ag<mailto: eblock@nde.ag>> wrote:
Can you add some more details as requested by Frank? Which mgr modules are enabled? What's the current 'ceph -s' output?
Is autoscaler running and doing stuff? Is balancer running and doing stuff? Is backfill going on? Is recovery going on? Is your ceph version affected by the "excessive logging to MON store" issue that was present starting with pacific but should have been addressed
Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto:zakhar@gmail.com :
We don't use CephFS at all and don't have RBD snapshots apart from some cloning for Openstack images.
The size of mon stores isn't an issue, it's < 600 MB. But it gets overwritten often causing lots of disk writes, and that is an issue for us.
/Z
On Wed, 11 Oct 2023 at 14:37, Eugen Block <eblock@nde.ag<mailto: eblock@nde.ag>> wrote:
Do you use many snapshots (rbd or cephfs)? That can cause a heavy monitor usage, we've seen large mon stores on customer clusters with rbd mirroring on snapshot basis. In a healthy cluster they have mon stores of around 2GB in size.
>> @Eugen: Was there not an option to limit logging to the MON store?
I don't recall at the moment, worth checking tough.
Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto: zakhar@gmail.com>>:
> Thank you, Frank. > > The cluster is healthy, operating normally, nothing unusual is going on. We > observe lots of writes by mon processes into mon rocksdb stores, > specifically: > > /var/lib/ceph/mon/ceph-cephXX/store.db: > 65M 3675511.sst > 65M 3675512.sst > 65M 3675513.sst > 65M 3675514.sst > 65M 3675515.sst > 65M 3675516.sst > 65M 3675517.sst > 65M 3675518.sst > 62M 3675519.sst > > The site of the files is not huge, but monitors rotate and write out these > files often, sometimes several times per minute, resulting in lots of data > written to disk. The writes coincide with "manual compaction" events logged > by the monitors, for example: > > debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1676] [default] [JOB 70854] Compacting 1@5 + > 9@6 files to L6, score -1.00 > debug 2023-10-11T11:10:10.483+0000 7f48a3a9b700 4 rocksdb: EVENT_LOG_v1 > {"time_micros": 1697022610487624, "job": 70854, "event": > "compaction_started", "compaction_reason": "ManualCompaction", "files_L5": > [3675543], "files_L6": [3675533, 3675534, 3675535, 3675536, 3675537, > 3675538, 3675539, 3675540, 3675541], "score": -1, "input_data_size": > 601117031} > debug 2023-10-11T11:10:10.619+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675544: 2015 keys, 67287115 bytes > debug 2023-10-11T11:10:10.763+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675545: 24343 keys, 67336225 bytes > debug 2023-10-11T11:10:10.899+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675546: 1196 keys, 67225813 bytes > debug 2023-10-11T11:10:11.035+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675547: 1049 keys, 67252678 bytes > debug 2023-10-11T11:10:11.167+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675548: 1081 keys, 67216638 bytes > debug 2023-10-11T11:10:11.303+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675549: 1196 keys, 67245376 bytes > debug 2023-10-11T11:10:12.023+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675550: 1195 keys, 67246813 bytes > debug 2023-10-11T11:10:13.059+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675551: 1205 keys, 67223302 bytes > debug 2023-10-11T11:10:13.903+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1349] [default] [JOB 70854] Generated table > #3675552: 1312 keys, 56416011 bytes > debug 2023-10-11T11:10:13.911+0000 7f48a3a9b700 4 rocksdb: > [compaction/compaction_job.cc:1415] [default] [JOB 70854] Compacted 1@5 + > 9@6 files to L6 => 594449971 bytes > debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log > Time 2023/10/11-11:10:13.920991) [compaction/compaction_job.cc:760] > [default] compacted to: base level 5 level multiplier 10.00 max bytes base > 268435456 files[0 0 0 0 0 0 9] max score 0.00, MB/sec: 175.8 rd, 173.9 wr, > level 6, files in(1, 9) out(9) MB in(0.3, 572.9) out(566.9), > read-write-amplify(3434.6) write-amplify(1707.7) OK, records in: 35108, > records dropped: 516 output_compression: NoCompression > debug 2023-10-11T11:10:13.915+0000 7f48a3a9b700 4 rocksdb: (Original Log > Time 2023/10/11-11:10:13.921010) EVENT_LOG_v1 {"time_micros": > 1697022613921002, "job": 70854, "event": "compaction_finished", > "compaction_time_micros": 3418822, "compaction_time_cpu_micros": 785454, > "output_level": 6, "num_output_files": 9, "total_output_size": 594449971, > "num_input_records": 35108, "num_output_records": 34592, > "num_subcompactions": 1, "output_compression": "NoCompression", > "num_single_delete_mismatches": 0, "num_single_delete_fallthrough": 0, > "lsm_state": [0, 0, 0, 0, 0, 0, 9]} > > The log even mentions the huge write multiplication. I wonder whether this > is normal and what can be done about it. > > /Z > > On Wed, 11 Oct 2023 at 13:55, Frank Schilder <frans@dtu.dk <mailto:frans@dtu.dk>> wrote: > >> I need to ask here: where exactly do you observe the hundreds of GB >> written per day? Are the mon logs huge? Is it the mon store? Is your >> cluster unhealthy? >> >> We have an octopus cluster with 1282 OSDs, 1650 ceph fs clients and about >> 800 librbd clients. Per week our mon logs are about 70M, the cluster logs >> about 120M , the audit logs about 70M and I see between 100-200Kb/s writes >> to the mon store. That's in the lower-digit GB range per day. Hundreds of >> GB per day sound completely over the top on a healthy cluster, unless you >> have MGR modules changing the OSD/cluster map continuously. >> >> Is autoscaler running and doing stuff? >> Is balancer running and doing stuff? >> Is backfill going on? >> Is recovery going on? >> Is your ceph version affected by the "excessive logging to MON store" >> issue that was present starting with pacific but should have been addressed >> by now? >> >> @Eugen: Was there not an option to limit logging to the MON store? >> >> For information to readers, we followed old recommendations from a Dell >> white paper for building a ceph cluster and have a 1TB Raid10 array on 6x >> write intensive SSDs for the MON stores. After 5 years we are below 10% >> wear. Average size of the MON store for a healthy cluster is 500M-1G, but >> we have seen this ballooning to 100+GB in degraded conditions. >> >> Best regards, >> ================= >> Frank Schilder >> AIT Risø Campus >> Bygning 109, rum S14 >> >> ________________________________________ >> From: Zakhar Kirpichenko <zakhar@gmail.com<mailto: zakhar@gmail.com>> >> Sent: Wednesday, October 11, 2023 12:00 PM >> To: Eugen Block >> Cc: ceph-users@ceph.io<mailto:ceph-users@ceph.io> >> Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes >> >> Thank you, Eugen. >> >> I'm interested specifically to find out whether the huge amount of data >> written by monitors is expected. It is eating through the endurance of our >> system drives, which were not specced for high DWPD/TBW, as this is not a >> documented requirement, and monitors produce hundreds of gigabytes of >> writes per day. I am looking for ways to reduce the amount of writes, if >> possible. >> >> /Z >> >> On Wed, 11 Oct 2023 at 12:41, Eugen Block <eblock@nde.ag<mailto: eblock@nde.ag>> wrote: >> >> > Hi, >> > >> > what you report is the expected behaviour, at least I see the same on >> > all clusters. I can't answer why the compaction is required that >> > often, but you can control the log level of the rocksdb output: >> > >> > ceph config set mon debug_rocksdb 1/5 (default is 4/5) >> > >> > This reduces the log entries and you wouldn't see the manual >> > compaction logs anymore. There are a couple more rocksdb options but I >> > probably wouldn't change too much, only if you know what you're doing. >> > Maybe Igor can comment if some other tuning makes sense here. >> > >> > Regards, >> > Eugen >> > >> > Zitat von Zakhar Kirpichenko <zakhar@gmail.com<mailto: zakhar@gmail.com>>: >> > >> > > Any input from anyone, please? >> > > >> > > On Tue, 10 Oct 2023 at 09:44, Zakhar Kirpichenko < zakhar@gmail.com<mailto:zakhar@gmail.com>> >> > wrote: >> > > >> > >> Any input from anyone, please? >> > >> >> > >> It's another thing that seems to be rather poorly documented: it's >> > unclear >> > >> what to expect, what 'normal' behavior should be, and what can be done >> > >> about the huge amount of writes by monitors. >> > >> >> > >> /Z >> > >> >> > >> On Mon, 9 Oct 2023 at 12:40, Zakhar Kirpichenko < zakhar@gmail.com<mailto:zakhar@gmail.com>> >> > wrote: >> > >> >> > >>> Hi, >> > >>> >> > >>> Monitors in our 16.2.14 cluster appear to quite often run "manual >> > >>> compaction" tasks: >> > >>> >> > >>> debug 2023-10-09T09:30:53.888+0000 7f48a329a700 4 rocksdb: >> > EVENT_LOG_v1 >> > >>> {"time_micros": 1696843853892760, "job": 64225, "event": >> > "flush_started", >> > >>> "num_memtables": 1, "num_entries": 715, "num_deletes": 251, >> > >>> "total_data_size": 3870352, "memory_usage": 3886744, "flush_reason": >> > >>> "Manual Compaction"} >> > >>> debug 2023-10-09T09:30:53.904+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original >> > Log >> > >>> Time 2023/10/09-09:30:53.910204) >> > [db_impl/db_impl_compaction_flush.cc:2516] >> > >>> [default] Manual compaction from level-0 to level-5 from 'paxos .. >> > 'paxos; >> > >>> will stop at (end) >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:30:53.908+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:30:53.908+0000 7f48a3a9b700 4 rocksdb: (Original >> > Log >> > >>> Time 2023/10/09-09:30:53.911004) >> > [db_impl/db_impl_compaction_flush.cc:2516] >> > >>> [default] Manual compaction from level-5 to level-6 from 'paxos .. >> > 'paxos; >> > >>> will stop at (end) >> > >>> debug 2023-10-09T09:32:08.956+0000 7f48a329a700 4 rocksdb: >> > EVENT_LOG_v1 >> > >>> {"time_micros": 1696843928961390, "job": 64228, "event": >> > "flush_started", >> > >>> "num_memtables": 1, "num_entries": 1580, "num_deletes": 502, >> > >>> "total_data_size": 8404605, "memory_usage": 8465840, "flush_reason": >> > >>> "Manual Compaction"} >> > >>> debug 2023-10-09T09:32:08.972+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original >> > Log >> > >>> Time 2023/10/09-09:32:08.977739) >> > [db_impl/db_impl_compaction_flush.cc:2516] >> > >>> [default] Manual compaction from level-0 to level-5 from 'logm .. >> > 'logm; >> > >>> will stop at (end) >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:08.976+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:08.976+0000 7f48a3a9b700 4 rocksdb: (Original >> > Log >> > >>> Time 2023/10/09-09:32:08.978512) >> > [db_impl/db_impl_compaction_flush.cc:2516] >> > >>> [default] Manual compaction from level-5 to level-6 from 'logm .. >> > 'logm; >> > >>> will stop at (end) >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:32:12.764+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:33:29.028+0000 7f48a329a700 4 rocksdb: >> > EVENT_LOG_v1 >> > >>> {"time_micros": 1696844009033151, "job": 64231, "event": >> > "flush_started", >> > >>> "num_memtables": 1, "num_entries": 1430, "num_deletes": 251, >> > >>> "total_data_size": 8975535, "memory_usage": 9035920, "flush_reason": >> > >>> "Manual Compaction"} >> > >>> debug 2023-10-09T09:33:29.044+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original >> > Log >> > >>> Time 2023/10/09-09:33:29.049585) >> > [db_impl/db_impl_compaction_flush.cc:2516] >> > >>> [default] Manual compaction from level-0 to level-5 from 'paxos .. >> > 'paxos; >> > >>> will stop at (end) >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:33:29.048+0000 7f4899286700 4 rocksdb: >> > >>> [db_impl/db_impl_compaction_flush.cc:1443] [default] Manual >> compaction >> > >>> starting >> > >>> debug 2023-10-09T09:33:29.048+0000 7f48a3a9b700 4 rocksdb: (Original >> > Log >> > >>> Time 2023/10/09-09:33:29.050355) >> > [db_impl/db_impl_compaction_flush.cc:2516] >> > >>> [default] Manual compaction from level-5 to level-6 from 'paxos .. >> > 'paxos; >> > >>> will stop at (end) >> > >>> >> > >>> I have removed a lot of interim log messages to save space. >> > >>> >> > >>> During each compaction the monitor process writes approximately >> 500-600 >> > >>> MB of data to disk over a short period of time. These writes add up >> to >> > tens >> > >>> of gigabytes per hour and hundreds of gigabytes per day. >> > >>> >> > >>> Monitor rocksdb and compaction options are default: >> > >>> >> > >>> "mon_compact_on_bootstrap": "false", >> > >>> "mon_compact_on_start": "false", >> > >>> "mon_compact_on_trim": "true", >> > >>> "mon_rocksdb_options": >> > >>> >> > >>
"write_buffer_size=33554432,compression=kNoCompression,level_compaction_dynamic_level_bytes=true",
>> > >>> >> > >>> Is this expected behavior? Is this something I can adjust in order to >> > >>> extend the system storage life? >> > >>> >> > >>> Best regards, >> > >>> Zakhar >> > >>> >> > >> >> > > _______________________________________________ >> > > ceph-users mailing list -- ceph-users@ceph.io<mailto: ceph-users@ceph.io> >> > > To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io> >> > >> > >> > _______________________________________________ >> > ceph-users mailing list -- ceph-users@ceph.io<mailto: ceph-users@ceph.io> >> > To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io> >> > >> _______________________________________________ >> ceph-users mailing list -- ceph-users@ceph.io<mailto: ceph-users@ceph.io> >> To unsubscribe send an email to ceph-users-leave@ceph.io<mailto: ceph-users-leave@ceph.io> >>
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all. /Z On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com> wrote:
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com <http://docs.ceph.com/> I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com <mailto:anthony.datri@gmail.com>> wrote:
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com <mailto:zakhar@gmail.com>> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/ Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community. /Z On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri <anthony.datri@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com> wrote:
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
The issue persists, although to a lesser extent. Any comments from the Ceph team please? /Z On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/
Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community.
/Z
On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri <anthony.datri@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com> wrote:
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well. Monitor w/o compression writes about 750 MB to disk in 5 minutes: 4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] Monitor with LZ4 compression writes about 1/4 of that over the same time period: 2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] This is caused by the apparent difference in store.db sizes. Mon store.db w/o compression: # ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst Mon store.db with compression: # ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors. /Z On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
The issue persists, although to a lesser extent. Any comments from the Ceph team please?
/Z
On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/
Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community.
/Z
On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri <anthony.datri@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com> wrote:
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
Hi Zakhar, I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values: paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250) This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do that in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well. Regards, Eugen https://tracker.ceph.com/issues/63229 Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well.
Monitor w/o compression writes about 750 MB to disk in 5 minutes:
4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
Monitor with LZ4 compression writes about 1/4 of that over the same time period:
2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
This is caused by the apparent difference in store.db sizes.
Mon store.db w/o compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst
Mon store.db with compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst
There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors.
/Z
On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
The issue persists, although to a lesser extent. Any comments from the Ceph team please?
/Z
On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/
Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community.
/Z
On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri <anthony.datri@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com> wrote:
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight! Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst to about half of the original size: -rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store: -rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable: Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes: -rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst # du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/ 2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes: -rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst # du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ 2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0] I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory this should result in lower but much faster compression. I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues. /Z On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
Hi Zakhar,
I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values:
paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250)
This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do that in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well.
Regards, Eugen
https://tracker.ceph.com/issues/63229
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well.
Monitor w/o compression writes about 750 MB to disk in 5 minutes:
4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
Monitor with LZ4 compression writes about 1/4 of that over the same time period:
2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
This is caused by the apparent difference in store.db sizes.
Mon store.db w/o compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst
Mon store.db with compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst
There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors.
/Z
On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
The issue persists, although to a lesser extent. Any comments from the Ceph team please?
/Z
On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/
Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community.
/Z
On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri <anthony.datri@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com
wrote:
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Zakhar, since its a bit beyond of the scope of basic, could you please post the complete ceph.conf config section for these changes for reference? Thanks! ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14 ________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 18, 2023 6:14 AM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight! Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB: 65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst to about half of the original size: -rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store: -rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable: Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes: -rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst # du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/ 2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes: -rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst # du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ 2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0] I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory this should result in lower but much faster compression. I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues. /Z On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
Hi Zakhar,
I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values:
paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250)
This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do that in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well.
Regards, Eugen
https://tracker.ceph.com/issues/63229
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well.
Monitor w/o compression writes about 750 MB to disk in 5 minutes:
4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
Monitor with LZ4 compression writes about 1/4 of that over the same time period:
2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
This is caused by the apparent difference in store.db sizes.
Mon store.db w/o compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst
Mon store.db with compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst
There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors.
/Z
On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
The issue persists, although to a lesser extent. Any comments from the Ceph team please?
/Z
On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/
Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community.
/Z
On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri <anthony.datri@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com
wrote:
cf. Mark's article I sent you re RocksDB tuning. I suspect that with Reef you would experience fewer writes. Universal compaction might also help, but in the end this SSD is a client SKU and really not suited for enterprise use. If you had the 1TB SKU you'd get much longer life, or you could change the overprovisioning on the ones you have.
On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
I would very much appreciate it if someone with a better understanding of monitor internals and use of RocksDB could please chip in.
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Frank, The only changes in ceph.conf are just the compression settings, most of the cluster configuration is in the monitor database thus my ceph.conf is rather short: --- [global] fsid = xxx mon_host = [list of mons] [mon.yyy] public network = a.b.c.d/e mon_rocksdb_options = "write_buffer_size=33554432,compression=kLZ4Compression,level_compaction_dynamic_level_bytes=true,bottommost_compression=kLZ4HCCompression" --- Note that my bottommost_compression choice is LZ4HC, whose compression is better than LZ4 at the expense of higher CPU usage. My nodes have lots of CPU to spare, so I went for LZ4HC for better space savings and a lower amount of writes. In general, I would recommend trying a faster and less intense compression first, LZ4 across the board is a good starting choice. /Z On Wed, 18 Oct 2023 at 12:02, Frank Schilder <frans@dtu.dk> wrote:
Hi Zakhar,
since its a bit beyond of the scope of basic, could you please post the complete ceph.conf config section for these changes for reference?
Thanks! ================= Frank Schilder AIT Risø Campus Bygning 109, rum S14
________________________________________ From: Zakhar Kirpichenko <zakhar@gmail.com> Sent: Wednesday, October 18, 2023 6:14 AM To: Eugen Block Cc: ceph-users@ceph.io Subject: [ceph-users] Re: Ceph 16.2.x mon compactions, disk writes
Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight!
Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB:
65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
to about half of the original size:
-rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst
Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store:
-rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst
Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable:
Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes:
-rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/ 2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0]
Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes:
-rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ 2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0]
I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory this should result in lower but much faster compression.
I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues.
/Z
On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
Hi Zakhar,
I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values:
paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250)
This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do that in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well.
Regards, Eugen
https://tracker.ceph.com/issues/63229
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well.
Monitor w/o compression writes about 750 MB to disk in 5 minutes:
4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
Monitor with LZ4 compression writes about 1/4 of that over the same time period:
2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
This is caused by the apparent difference in store.db sizes.
Mon store.db w/o compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst
Mon store.db with compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst
There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors.
/Z
On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
The issue persists, although to a lesser extent. Any comments from the Ceph team please?
/Z
On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/
Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community.
/Z
On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri < anthony.datri@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri < anthony.datri@gmail.com
wrote:
> cf. Mark's article I sent you re RocksDB tuning. I suspect that with > Reef you would experience fewer writes. Universal compaction might also > help, but in the end this SSD is a client SKU and really not suited for > enterprise use. If you had the 1TB SKU you'd get much longer > life, or you > could change the overprovisioning on the ones you have. > > On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote: > > I would very much appreciate it if someone with a better understanding > of > monitor internals and use of RocksDB could please chip in. > > >
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Zakhar, hello List, I just wanted to follow up on this and ask a few quesitions: Did you noticed any downsides with your compression settings so far? Do you have all mons now on compression? Did release updates go through without issues? Do you know if this works also with reef (we see massive writes as well there)? Can you briefly tabulate the commands you used to persistently set the compression options? Thanks so much, Dietmar On 10/18/23 06:14, Zakhar Kirpichenko wrote:
Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight!
Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB:
65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
to about half of the original size:
-rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst
Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store:
-rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst
Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable:
Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes:
-rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/ 2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0]
Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes:
-rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ 2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0]
I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory this should result in lower but much faster compression.
I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues.
/Z
On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
Hi Zakhar,
I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values:
paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250)
This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do that in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well.
Regards, Eugen
https://tracker.ceph.com/issues/63229
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well.
Monitor w/o compression writes about 750 MB to disk in 5 minutes:
4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
Monitor with LZ4 compression writes about 1/4 of that over the same time period:
2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
This is caused by the apparent difference in store.db sizes.
Mon store.db w/o compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst
Mon store.db with compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst
There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors.
/Z
On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
The issue persists, although to a lesser extent. Any comments from the Ceph team please?
/Z
On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/
Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community.
/Z
On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri <anthony.datri@gmail.com> wrote:
Some of it is transferable to RocksDB on mons nonetheless.
but their specs exceed Ceph hardware recommendations by a good margin
Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
Thank you, Anthony. As I explained to you earlier, the article you had sent is about RocksDB tuning for Bluestore OSDs, while the issue at hand is not with OSDs but rather monitors and their RocksDB store. Indeed, the drives are not enterprise-grade, but their specs exceed Ceph hardware recommendations by a good margin, they're being used as boot drives only and aren't supposed to be written to continuously at high rates - which is what unfortunately is happening. I am trying to determine why it is happening and how the issue can be alleviated or resolved, unfortunately monitor RocksDB usage and tunables appear to be not documented at all.
/Z
On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com
wrote:
> cf. Mark's article I sent you re RocksDB tuning. I suspect that with > Reef you would experience fewer writes. Universal compaction might also > help, but in the end this SSD is a client SKU and really not suited for > enterprise use. If you had the 1TB SKU you'd get much longer > life, or you > could change the overprovisioning on the ones you have. > > On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote: > > I would very much appreciate it if someone with a better understanding > of > monitor internals and use of RocksDB could please chip in. > > >
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
Hi,
Did you noticed any downsides with your compression settings so far?
None, at least on our systems. Except the part that I haven't found a way to make the settings persist.
Do you have all mons now on compression?
I have 3 out of 5 monitors with compression and 2 without it. The 2 monitors with uncompressed RocksDB have much larger disks which do not suffer from writes as much as the other 3. I keep them uncompressed "just in case", i.e. for the unlikely event if the 3 monitors with compressed RocksDB fail or have any issues specifically because of the compression. I have to say that this hasn't happened yet, and this precaution may be unnecessary.
Did release updates go through without issues?
Do you know if this works also with reef (we see massive writes as well
In our case, container updates overwrite the monitors' configurations and reset RocksDB options, thus each updated monitor runs with no RocksDB compression until it is added back manually. Other than that, I have not encountered any issues related to compression during the updates. there)? Unfortunately, I can't comment on Reef as we're still using Pacific. /Z On Tue, 16 Apr 2024 at 18:08, Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
Hi Zakhar, hello List,
I just wanted to follow up on this and ask a few quesitions:
Did you noticed any downsides with your compression settings so far? Do you have all mons now on compression? Did release updates go through without issues? Do you know if this works also with reef (we see massive writes as well there)?
Can you briefly tabulate the commands you used to persistently set the compression options?
Thanks so much,
Dietmar
Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight!
Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB:
65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
to about half of the original size:
-rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst
Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store:
-rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst
Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable:
Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes:
-rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/ 2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0]
Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes:
-rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ 2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0]
I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory this should result in lower but much faster compression.
I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues.
/Z
On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
Hi Zakhar,
I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values:
paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250)
This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do that in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well.
Regards, Eugen
https://tracker.ceph.com/issues/63229
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well.
Monitor w/o compression writes about 750 MB to disk in 5 minutes:
4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
Monitor with LZ4 compression writes about 1/4 of that over the same time period:
2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
This is caused by the apparent difference in store.db sizes.
Mon store.db w/o compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst
Mon store.db with compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst
There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors.
/Z
On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
The issue persists, although to a lesser extent. Any comments from the Ceph team please?
/Z
On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
> Some of it is transferable to RocksDB on mons nonetheless.
Please point me to relevant Ceph documentation, i.e. a description of how various Ceph monitor and RocksDB tunables affect the operations of monitors, I'll gladly look into it.
> Please point me to such recommendations, if they're on docs.ceph.com I'll get them updated.
This are the recommendations we used when we built our Pacific cluster: https://docs.ceph.com/en/pacific/start/hardware-recommendations/
Our drives are 4x times larger than recommended by this guide. The drives are rated for < 0.5 DWPD, which is more than sufficient for boot drives and storage of rarely modified files. It is not documented or suggested anywhere that monitor processes write several hundred gigabytes of data per day, exceeding the amount of data written by OSDs. Which is why I am not convinced that what we're observing is expected behavior, but it's not easy to get a definitive answer from the Ceph community.
/Z
On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri < anthony.datri@gmail.com> wrote:
> Some of it is transferable to RocksDB on mons nonetheless. > > but their specs exceed Ceph hardware recommendations by a good margin > > > Please point me to such recommendations, if they're on docs.ceph.com I'll > get them updated. > > On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote: > > Thank you, Anthony. As I explained to you earlier, the article you had > sent is about RocksDB tuning for Bluestore OSDs, while the issue > at hand is > not with OSDs but rather monitors and their RocksDB store. Indeed,
On 10/18/23 06:14, Zakhar Kirpichenko wrote: the
> drives are not enterprise-grade, but their specs exceed Ceph hardware > recommendations by a good margin, they're being used as boot drives only > and aren't supposed to be written to continuously at high rates - which is > what unfortunately is happening. I am trying to determine why it is > happening and how the issue can be alleviated or resolved, unfortunately > monitor RocksDB usage and tunables appear to be not documented at all. > > /Z > > On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri < anthony.datri@gmail.com
> wrote: > >> cf. Mark's article I sent you re RocksDB tuning. I suspect that with >> Reef you would experience fewer writes. Universal compaction might also >> help, but in the end this SSD is a client SKU and really not suited for >> enterprise use. If you had the 1TB SKU you'd get much longer >> life, or you >> could change the overprovisioning on the ones you have. >> >> On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote: >> >> I would very much appreciate it if someone with a better understanding >> of >> monitor internals and use of RocksDB could please chip in. >> >> >> >
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
You can use the extra container arguments I pointed out a few months ago. Those work in my test clusters, although I haven’t enabled that in production yet. But it shouldn’t make a difference if it’s a test cluster or not. 😉 Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Hi,
Did you noticed any downsides with your compression settings so far?
None, at least on our systems. Except the part that I haven't found a way to make the settings persist.
Do you have all mons now on compression?
I have 3 out of 5 monitors with compression and 2 without it. The 2 monitors with uncompressed RocksDB have much larger disks which do not suffer from writes as much as the other 3. I keep them uncompressed "just in case", i.e. for the unlikely event if the 3 monitors with compressed RocksDB fail or have any issues specifically because of the compression. I have to say that this hasn't happened yet, and this precaution may be unnecessary.
Did release updates go through without issues?
In our case, container updates overwrite the monitors' configurations and reset RocksDB options, thus each updated monitor runs with no RocksDB compression until it is added back manually. Other than that, I have not encountered any issues related to compression during the updates.
Do you know if this works also with reef (we see massive writes as well there)?
Unfortunately, I can't comment on Reef as we're still using Pacific.
/Z
On Tue, 16 Apr 2024 at 18:08, Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
Hi Zakhar, hello List,
I just wanted to follow up on this and ask a few quesitions:
Did you noticed any downsides with your compression settings so far? Do you have all mons now on compression? Did release updates go through without issues? Do you know if this works also with reef (we see massive writes as well there)?
Can you briefly tabulate the commands you used to persistently set the compression options?
Thanks so much,
Dietmar
Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight!
Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB:
65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
to about half of the original size:
-rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst
Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store:
-rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst
Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable:
Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes:
-rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/ 2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0]
Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes:
-rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ 2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0]
I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory this should result in lower but much faster compression.
I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues.
/Z
On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
Hi Zakhar,
I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values:
paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250)
This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do that in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well.
Regards, Eugen
https://tracker.ceph.com/issues/63229
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well.
Monitor w/o compression writes about 750 MB to disk in 5 minutes:
4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
Monitor with LZ4 compression writes about 1/4 of that over the same time period:
2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
This is caused by the apparent difference in store.db sizes.
Mon store.db w/o compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst
Mon store.db with compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst
There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors.
/Z
On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
The issue persists, although to a lesser extent. Any comments from the Ceph team please?
/Z
On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
>> Some of it is transferable to RocksDB on mons nonetheless. > > Please point me to relevant Ceph documentation, i.e. a description of how > various Ceph monitor and RocksDB tunables affect the operations of > monitors, I'll gladly look into it. > >> Please point me to such recommendations, if they're on docs.ceph.com I'll > get them updated. > > This are the recommendations we used when we built our Pacific cluster: > https://docs.ceph.com/en/pacific/start/hardware-recommendations/ > > Our drives are 4x times larger than recommended by this guide. The drives > are rated for < 0.5 DWPD, which is more than sufficient for boot drives and > storage of rarely modified files. It is not documented or suggested > anywhere that monitor processes write several hundred gigabytes of data per > day, exceeding the amount of data written by OSDs. Which is why I am not > convinced that what we're observing is expected behavior, but it's not easy > to get a definitive answer from the Ceph community. > > /Z > > On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri < anthony.datri@gmail.com> > wrote: > >> Some of it is transferable to RocksDB on mons nonetheless. >> >> but their specs exceed Ceph hardware recommendations by a good margin >> >> >> Please point me to such recommendations, if they're on docs.ceph.com I'll >> get them updated. >> >> On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote: >> >> Thank you, Anthony. As I explained to you earlier, the article you had >> sent is about RocksDB tuning for Bluestore OSDs, while the issue >> at hand is >> not with OSDs but rather monitors and their RocksDB store. Indeed,
On 10/18/23 06:14, Zakhar Kirpichenko wrote: the
>> drives are not enterprise-grade, but their specs exceed Ceph hardware >> recommendations by a good margin, they're being used as boot drives only >> and aren't supposed to be written to continuously at high rates - which is >> what unfortunately is happening. I am trying to determine why it is >> happening and how the issue can be alleviated or resolved, unfortunately >> monitor RocksDB usage and tunables appear to be not documented at all. >> >> /Z >> >> On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri < anthony.datri@gmail.com
>> wrote: >> >>> cf. Mark's article I sent you re RocksDB tuning. I suspect that with >>> Reef you would experience fewer writes. Universal compaction might also >>> help, but in the end this SSD is a client SKU and really not suited for >>> enterprise use. If you had the 1TB SKU you'd get much longer >>> life, or you >>> could change the overprovisioning on the ones you have. >>> >>> On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote: >>> >>> I would very much appreciate it if someone with a better understanding >>> of >>> monitor internals and use of RocksDB could please chip in. >>> >>> >>> >>
ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Sorry, I meant extra-entrypoint-arguments: https://www.spinics.net/lists/ceph-users/msg79251.html Zitat von Eugen Block <eblock@nde.ag>:
You can use the extra container arguments I pointed out a few months ago. Those work in my test clusters, although I haven’t enabled that in production yet. But it shouldn’t make a difference if it’s a test cluster or not. 😉
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Hi,
Did you noticed any downsides with your compression settings so far?
None, at least on our systems. Except the part that I haven't found a way to make the settings persist.
Do you have all mons now on compression?
I have 3 out of 5 monitors with compression and 2 without it. The 2 monitors with uncompressed RocksDB have much larger disks which do not suffer from writes as much as the other 3. I keep them uncompressed "just in case", i.e. for the unlikely event if the 3 monitors with compressed RocksDB fail or have any issues specifically because of the compression. I have to say that this hasn't happened yet, and this precaution may be unnecessary.
Did release updates go through without issues?
In our case, container updates overwrite the monitors' configurations and reset RocksDB options, thus each updated monitor runs with no RocksDB compression until it is added back manually. Other than that, I have not encountered any issues related to compression during the updates.
Do you know if this works also with reef (we see massive writes as well there)?
Unfortunately, I can't comment on Reef as we're still using Pacific.
/Z
On Tue, 16 Apr 2024 at 18:08, Dietmar Rieder <dietmar.rieder@i-med.ac.at> wrote:
Hi Zakhar, hello List,
I just wanted to follow up on this and ask a few quesitions:
Did you noticed any downsides with your compression settings so far? Do you have all mons now on compression? Did release updates go through without issues? Do you know if this works also with reef (we see massive writes as well there)?
Can you briefly tabulate the commands you used to persistently set the compression options?
Thanks so much,
Dietmar
Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight!
Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB:
65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
to about half of the original size:
-rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst
Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store:
-rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst
Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable:
Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes:
-rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/ 2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0]
Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes:
-rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst
# du -hs /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/; iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ 2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0]
I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory this should result in lower but much faster compression.
I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues.
/Z
On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
Hi Zakhar,
I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values:
paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250)
This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do that in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well.
Regards, Eugen
https://tracker.ceph.com/issues/63229
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
With the help of community members, I managed to enable RocksDB compression for a test monitor, and it seems to be working well.
Monitor w/o compression writes about 750 MB to disk in 5 minutes:
4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
Monitor with LZ4 compression writes about 1/4 of that over the same time period:
2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0]
This is caused by the apparent difference in store.db sizes.
Mon store.db w/o compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db total 257196 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Aug 31 05:22 .. -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst
Mon store.db with compression:
# ls -al /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db total 91188 drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . drwx------ 3 167 167 4096 Oct 16 13:35 .. -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst
There are no apparent downsides thus far. If everything works well, I will try adding compression to other monitors.
/Z
On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote:
> The issue persists, although to a lesser extent. Any comments from the > Ceph team please? > > /Z > > On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com> wrote: > >>> Some of it is transferable to RocksDB on mons nonetheless. >> >> Please point me to relevant Ceph documentation, i.e. a description of how >> various Ceph monitor and RocksDB tunables affect the operations of >> monitors, I'll gladly look into it. >> >>> Please point me to such recommendations, if they're on docs.ceph.com I'll >> get them updated. >> >> This are the recommendations we used when we built our Pacific cluster: >> https://docs.ceph.com/en/pacific/start/hardware-recommendations/ >> >> Our drives are 4x times larger than recommended by this guide. The drives >> are rated for < 0.5 DWPD, which is more than sufficient for boot drives and >> storage of rarely modified files. It is not documented or suggested >> anywhere that monitor processes write several hundred gigabytes of data per >> day, exceeding the amount of data written by OSDs. Which is why I am not >> convinced that what we're observing is expected behavior, but it's not easy >> to get a definitive answer from the Ceph community. >> >> /Z >> >> On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri < anthony.datri@gmail.com> >> wrote: >> >>> Some of it is transferable to RocksDB on mons nonetheless. >>> >>> but their specs exceed Ceph hardware recommendations by a good margin >>> >>> >>> Please point me to such recommendations, if they're on docs.ceph.com I'll >>> get them updated. >>> >>> On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote: >>> >>> Thank you, Anthony. As I explained to you earlier, the article you had >>> sent is about RocksDB tuning for Bluestore OSDs, while the issue >>> at hand is >>> not with OSDs but rather monitors and their RocksDB store. Indeed,
On 10/18/23 06:14, Zakhar Kirpichenko wrote: the
>>> drives are not enterprise-grade, but their specs exceed Ceph hardware >>> recommendations by a good margin, they're being used as boot drives only >>> and aren't supposed to be written to continuously at high rates - which is >>> what unfortunately is happening. I am trying to determine why it is >>> happening and how the issue can be alleviated or resolved, unfortunately >>> monitor RocksDB usage and tunables appear to be not documented at all. >>> >>> /Z >>> >>> On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri < anthony.datri@gmail.com
>>> wrote: >>> >>>> cf. Mark's article I sent you re RocksDB tuning. I suspect that with >>>> Reef you would experience fewer writes. Universal compaction might also >>>> help, but in the end this SSD is a client SKU and really not suited for >>>> enterprise use. If you had the 1TB SKU you'd get much longer >>>> life, or you >>>> could change the overprovisioning on the ones you have. >>>> >>>> On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com> wrote: >>>> >>>> I would very much appreciate it if someone with a better understanding >>>> of >>>> monitor internals and use of RocksDB could please chip in. >>>> >>>> >>>> >>> _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I remember that I found the part which said "if something goes wrong, monitors will fail" rather discouraging :-) /Z On Tue, 16 Apr 2024 at 18:59, Eugen Block <eblock@nde.ag> wrote:
Sorry, I meant extra-entrypoint-arguments:
https://www.spinics.net/lists/ceph-users/msg79251.html
Zitat von Eugen Block <eblock@nde.ag>:
You can use the extra container arguments I pointed out a few months ago. Those work in my test clusters, although I haven’t enabled that in production yet. But it shouldn’t make a difference if it’s a test cluster or not. 😉
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Hi,
Did you noticed any downsides with your compression settings so far?
None, at least on our systems. Except the part that I haven't found a way to make the settings persist.
Do you have all mons now on compression?
I have 3 out of 5 monitors with compression and 2 without it. The 2 monitors with uncompressed RocksDB have much larger disks which do not suffer from writes as much as the other 3. I keep them uncompressed "just in case", i.e. for the unlikely event if the 3 monitors with compressed RocksDB fail or have any issues specifically because of the compression. I have to say that this hasn't happened yet, and this precaution may be unnecessary.
Did release updates go through without issues?
In our case, container updates overwrite the monitors' configurations and reset RocksDB options, thus each updated monitor runs with no RocksDB compression until it is added back manually. Other than that, I have not encountered any issues related to compression during the updates.
Do you know if this works also with reef (we see massive writes as well there)?
Unfortunately, I can't comment on Reef as we're still using Pacific.
/Z
On Tue, 16 Apr 2024 at 18:08, Dietmar Rieder < dietmar.rieder@i-med.ac.at> wrote:
Hi Zakhar, hello List,
I just wanted to follow up on this and ask a few quesitions:
Did you noticed any downsides with your compression settings so far? Do you have all mons now on compression? Did release updates go through without issues? Do you know if this works also with reef (we see massive writes as well there)?
Can you briefly tabulate the commands you used to persistently set the compression options?
Thanks so much,
Dietmar
On 10/18/23 06:14, Zakhar Kirpichenko wrote:
Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight!
Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB:
65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
to about half of the original size:
-rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst
Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store:
-rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst
Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable:
Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes:
-rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst
# du -hs
/var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/;
iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M
/var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/
2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0]
Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes:
-rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst
# du -hs
/var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/;
iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M
2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0]
I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory
should result in lower but much faster compression.
I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues.
/Z
On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
Hi Zakhar,
I took a closer look into what the MONs really do (again with Mykola's help) and why manual compaction is triggered so frequently. With debug_paxos=20 I noticed that paxosservice and paxos triggered manual compactions. So I played with these values:
paxos_service_trim_max = 1000 (default 500) paxos_service_trim_min = 500 (default 250) paxos_trim_max = 1000 (default 500) paxos_trim_min = 500 (default 250)
This reduced the amount of writes by a factor of 3 or 4, the iotop values are fluctuating a bit, of course. As Mykola suggested I created a tracker issue [1] to increase the default values since they don't seem suitable for a production environment. Although I don't have tested that in production yet I'll ask one of our customers to do
in their secondary cluster (for rbd mirroring) where they also suffer from large mon stores and heavy writes to the mon store. Your findings with the compaction were quite helpful as well, we'll test that as well. Igor mentioned that the default bluestore_rocksdb config for OSDs will enable compression because of positive test results. If we can confirm that compression works well for MONs too, compression could be enabled by default as well.
Regards, Eugen
https://tracker.ceph.com/issues/63229
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
> With the help of community members, I managed to enable RocksDB compression > for a test monitor, and it seems to be working well. > > Monitor w/o compression writes about 750 MB to disk in 5 minutes: > > 4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n > mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [rocksdb:low0] > > Monitor with LZ4 compression writes about 1/4 of that over the same time > period: > > 2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n > mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [rocksdb:low0] > > This is caused by the apparent difference in store.db sizes. > > Mon store.db w/o compression: > > # ls -al > /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db > total 257196 > drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . > drwx------ 3 167 167 4096 Aug 31 05:22 .. > -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log > -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst > -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst > -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst > > Mon store.db with compression: > > # ls -al > /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db > total 91188 > drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . > drwx------ 3 167 167 4096 Oct 16 13:35 .. > -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log > -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst > > There are no apparent downsides thus far. If everything works well, I will > try adding compression to other monitors. > > /Z > > On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> wrote: > >> The issue persists, although to a lesser extent. Any comments from
/var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ this that the
>> Ceph team please? >> >> /Z >> >> On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com
wrote: >> >>>> Some of it is transferable to RocksDB on mons nonetheless. >>> >>> Please point me to relevant Ceph documentation, i.e. a description of how >>> various Ceph monitor and RocksDB tunables affect the operations of >>> monitors, I'll gladly look into it. >>> >>>> Please point me to such recommendations, if they're on docs.ceph.com I'll >>> get them updated. >>> >>> This are the recommendations we used when we built our Pacific cluster: >>> https://docs.ceph.com/en/pacific/start/hardware-recommendations/ >>> >>> Our drives are 4x times larger than recommended by this guide. The drives >>> are rated for < 0.5 DWPD, which is more than sufficient for boot drives and >>> storage of rarely modified files. It is not documented or suggested >>> anywhere that monitor processes write several hundred gigabytes of data per >>> day, exceeding the amount of data written by OSDs. Which is why I am not >>> convinced that what we're observing is expected behavior, but it's not easy >>> to get a definitive answer from the Ceph community. >>> >>> /Z >>> >>> On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri < anthony.datri@gmail.com> >>> wrote: >>> >>>> Some of it is transferable to RocksDB on mons nonetheless. >>>> >>>> but their specs exceed Ceph hardware recommendations by a good margin >>>> >>>> >>>> Please point me to such recommendations, if they're on docs.ceph.com I'll >>>> get them updated. >>>> >>>> On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> wrote: >>>> >>>> Thank you, Anthony. As I explained to you earlier, the article you had >>>> sent is about RocksDB tuning for Bluestore OSDs, while the issue >>>> at hand is >>>> not with OSDs but rather monitors and their RocksDB store. Indeed, the >>>> drives are not enterprise-grade, but their specs exceed Ceph hardware >>>> recommendations by a good margin, they're being used as boot drives only >>>> and aren't supposed to be written to continuously at high rates - which is >>>> what unfortunately is happening. I am trying to determine why it is >>>> happening and how the issue can be alleviated or resolved, unfortunately >>>> monitor RocksDB usage and tunables appear to be not documented at all. >>>> >>>> /Z >>>> >>>> On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri < anthony.datri@gmail.com > >>>> wrote: >>>> >>>>> cf. Mark's article I sent you re RocksDB tuning. I suspect that with >>>>> Reef you would experience fewer writes. Universal compaction might also >>>>> help, but in the end this SSD is a client SKU and really not suited for >>>>> enterprise use. If you had the 1TB SKU you'd get much longer >>>>> life, or you >>>>> could change the overprovisioning on the ones you have. >>>>> >>>>> On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com
wrote: >>>>> >>>>> I would very much appreciate it if someone with a better understanding >>>>> of >>>>> monitor internals and use of RocksDB could please chip in. >>>>> >>>>> >>>>> >>>> > _______________________________________________ > ceph-users mailing list -- ceph-users@ceph.io > To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
I understand, I just wanted to point out to be careful (as you always should be with critical services like MONs). If you apply it to one mon first you will see if it works or not. Then apply it to the other ones. Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
I remember that I found the part which said "if something goes wrong, monitors will fail" rather discouraging :-)
/Z
On Tue, 16 Apr 2024 at 18:59, Eugen Block <eblock@nde.ag> wrote:
Sorry, I meant extra-entrypoint-arguments:
https://www.spinics.net/lists/ceph-users/msg79251.html
Zitat von Eugen Block <eblock@nde.ag>:
You can use the extra container arguments I pointed out a few months ago. Those work in my test clusters, although I haven’t enabled that in production yet. But it shouldn’t make a difference if it’s a test cluster or not. 😉
Zitat von Zakhar Kirpichenko <zakhar@gmail.com>:
Hi,
Did you noticed any downsides with your compression settings so far?
None, at least on our systems. Except the part that I haven't found a way to make the settings persist.
Do you have all mons now on compression?
I have 3 out of 5 monitors with compression and 2 without it. The 2 monitors with uncompressed RocksDB have much larger disks which do not suffer from writes as much as the other 3. I keep them uncompressed "just in case", i.e. for the unlikely event if the 3 monitors with compressed RocksDB fail or have any issues specifically because of the compression. I have to say that this hasn't happened yet, and this precaution may be unnecessary.
Did release updates go through without issues?
In our case, container updates overwrite the monitors' configurations and reset RocksDB options, thus each updated monitor runs with no RocksDB compression until it is added back manually. Other than that, I have not encountered any issues related to compression during the updates.
Do you know if this works also with reef (we see massive writes as well there)?
Unfortunately, I can't comment on Reef as we're still using Pacific.
/Z
On Tue, 16 Apr 2024 at 18:08, Dietmar Rieder < dietmar.rieder@i-med.ac.at> wrote:
Hi Zakhar, hello List,
I just wanted to follow up on this and ask a few quesitions:
Did you noticed any downsides with your compression settings so far? Do you have all mons now on compression? Did release updates go through without issues? Do you know if this works also with reef (we see massive writes as well there)?
Can you briefly tabulate the commands you used to persistently set the compression options?
Thanks so much,
Dietmar
On 10/18/23 06:14, Zakhar Kirpichenko wrote:
Many thanks for this, Eugen! I very much appreciate yours and Mykola's efforts and insight!
Another thing I noticed was a reduction of RocksDB store after the reduction of the total PG number by 30%, from 590-600 MB:
65M 3675511.sst 65M 3675512.sst 65M 3675513.sst 65M 3675514.sst 65M 3675515.sst 65M 3675516.sst 65M 3675517.sst 65M 3675518.sst 62M 3675519.sst
to about half of the original size:
-rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst
Then when I restarted the monitors one by one before adding compression, RocksDB store reduced even further. I am not sure why and what exactly got automatically removed from the store:
-rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst
Then I have enabled LZ4 and LZ4HC compression in our small production cluster (6 nodes, 96 OSDs) on 3 out of 5 monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. I specifically went for LZ4 and LZ4HC because of the balance between compression/decompression speed and impact on CPU usage. The compression doesn't seem to affect the cluster in any negative way, the 3 monitors with compression are operating normally. The effect of the compression on RocksDB store size and disk writes is quite noticeable:
Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB writes over 5 minutes:
-rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst
# du -hs
/var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/;
iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 155M
/var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/
2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0]
Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of writes over 5 minutes:
-rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst
# du -hs
/var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/;
iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon 60M
2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:low0] 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [ms_dispatch] 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [rocksdb:high0] 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false --default-log-to-stderr=true --default-log-stderr-prefix=debug --default-mon-cluster-log-to-file=false --default-mon-cluster-log-to-stderr=true [msgr-worker-0]
I haven't noticed a major CPU impact. Unfortunately I didn't specifically measure CPU time for monitors and , but overall the CPU impact of monitor store compression on our systems isn't noticeable. This may be different for larger clusters with larger RocksDB datasets, then perhaps compression=kLZ4Compression can be enabled by defualt and bottommost_compression=kLZ4HCCompression can be optional, in theory
should result in lower but much faster compression.
I hope this helps. My plan is to keep the monitors with the current settings, i.e. 3 with compression + 2 without compression, until the next minor release of Pacific to see whether the monitors with compressed RocksDB store can be upgraded without issues.
/Z
On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag> wrote:
> Hi Zakhar, > > I took a closer look into what the MONs really do (again with Mykola's > help) and why manual compaction is triggered so frequently. With > debug_paxos=20 I noticed that paxosservice and paxos triggered manual > compactions. So I played with these values: > > paxos_service_trim_max = 1000 (default 500) > paxos_service_trim_min = 500 (default 250) > paxos_trim_max = 1000 (default 500) > paxos_trim_min = 500 (default 250) > > This reduced the amount of writes by a factor of 3 or 4, the iotop > values are fluctuating a bit, of course. As Mykola suggested I created > a tracker issue [1] to increase the default values since they don't > seem suitable for a production environment. Although I don't have > tested that in production yet I'll ask one of our customers to do
> in their secondary cluster (for rbd mirroring) where they also suffer > from large mon stores and heavy writes to the mon store. Your findings > with the compaction were quite helpful as well, we'll test that as well. > Igor mentioned that the default bluestore_rocksdb config for OSDs will > enable compression because of positive test results. If we can confirm > that compression works well for MONs too, compression could be enabled > by default as well. > > Regards, > Eugen > > https://tracker.ceph.com/issues/63229 > > Zitat von Zakhar Kirpichenko <zakhar@gmail.com>: > >> With the help of community members, I managed to enable RocksDB > compression >> for a test monitor, and it seems to be working well. >> >> Monitor w/o compression writes about 750 MB to disk in 5 minutes: >> >> 4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n >> mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false >> --default-log-to-stderr=true --default-log-stderr-prefix=debug >> --default-mon-cluster-log-to-file=false >> --default-mon-cluster-log-to-stderr=true [rocksdb:low0] >> >> Monitor with LZ4 compression writes about 1/4 of that over the same time >> period: >> >> 2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n >> mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false >> --default-log-to-stderr=true --default-log-stderr-prefix=debug >> --default-mon-cluster-log-to-file=false >> --default-mon-cluster-log-to-stderr=true [rocksdb:low0] >> >> This is caused by the apparent difference in store.db sizes. >> >> Mon store.db w/o compression: >> >> # ls -al >> /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db >> total 257196 >> drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . >> drwx------ 3 167 167 4096 Aug 31 05:22 .. >> -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log >> -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst >> -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst >> -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst >> >> Mon store.db with compression: >> >> # ls -al >> /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db >> total 91188 >> drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . >> drwx------ 3 167 167 4096 Oct 16 13:35 .. >> -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log >> -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst >> >> There are no apparent downsides thus far. If everything works well, I > will >> try adding compression to other monitors. >> >> /Z >> >> On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com> > wrote: >> >>> The issue persists, although to a lesser extent. Any comments from
/var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ this that the
>>> Ceph team please? >>> >>> /Z >>> >>> On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com
> wrote: >>> >>>>> Some of it is transferable to RocksDB on mons nonetheless. >>>> >>>> Please point me to relevant Ceph documentation, i.e. a description of > how >>>> various Ceph monitor and RocksDB tunables affect the operations of >>>> monitors, I'll gladly look into it. >>>> >>>>> Please point me to such recommendations, if they're on docs.ceph.com > I'll >>>> get them updated. >>>> >>>> This are the recommendations we used when we built our Pacific cluster: >>>> https://docs.ceph.com/en/pacific/start/hardware-recommendations/ >>>> >>>> Our drives are 4x times larger than recommended by this guide. The > drives >>>> are rated for < 0.5 DWPD, which is more than sufficient for boot > drives and >>>> storage of rarely modified files. It is not documented or suggested >>>> anywhere that monitor processes write several hundred gigabytes of > data per >>>> day, exceeding the amount of data written by OSDs. Which is why I am > not >>>> convinced that what we're observing is expected behavior, but it's not > easy >>>> to get a definitive answer from the Ceph community. >>>> >>>> /Z >>>> >>>> On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri < anthony.datri@gmail.com> >>>> wrote: >>>> >>>>> Some of it is transferable to RocksDB on mons nonetheless. >>>>> >>>>> but their specs exceed Ceph hardware recommendations by a good margin >>>>> >>>>> >>>>> Please point me to such recommendations, if they're on docs.ceph.com > I'll >>>>> get them updated. >>>>> >>>>> On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com> > wrote: >>>>> >>>>> Thank you, Anthony. As I explained to you earlier, the article you had >>>>> sent is about RocksDB tuning for Bluestore OSDs, while the issue >>>>> at hand is >>>>> not with OSDs but rather monitors and their RocksDB store. Indeed, the >>>>> drives are not enterprise-grade, but their specs exceed Ceph hardware >>>>> recommendations by a good margin, they're being used as boot drives > only >>>>> and aren't supposed to be written to continuously at high rates - > which is >>>>> what unfortunately is happening. I am trying to determine why it is >>>>> happening and how the issue can be alleviated or resolved, > unfortunately >>>>> monitor RocksDB usage and tunables appear to be not documented at all. >>>>> >>>>> /Z >>>>> >>>>> On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri < anthony.datri@gmail.com >> >>>>> wrote: >>>>> >>>>>> cf. Mark's article I sent you re RocksDB tuning. I suspect that with >>>>>> Reef you would experience fewer writes. Universal compaction might > also >>>>>> help, but in the end this SSD is a client SKU and really not suited > for >>>>>> enterprise use. If you had the 1TB SKU you'd get much longer >>>>>> life, or you >>>>>> could change the overprovisioning on the ones you have. >>>>>> >>>>>> On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com
> wrote: >>>>>> >>>>>> I would very much appreciate it if someone with a better > understanding >>>>>> of >>>>>> monitor internals and use of RocksDB could please chip in. >>>>>> >>>>>> >>>>>> >>>>> >> _______________________________________________ >> ceph-users mailing list -- ceph-users@ceph.io >> To unsubscribe send an email to ceph-users-leave@ceph.io > > > _______________________________________________ > ceph-users mailing list -- ceph-users@ceph.io > To unsubscribe send an email to ceph-users-leave@ceph.io > _______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
-- _________________________________________________________ D i e t m a r R i e d e r Innsbruck Medical University Biocenter - Institute of Bioinformatics Innrain 80, 6020 Innsbruck Phone: +43 512 9003 71402 | Mobile: +43 676 8716 72402 Email: dietmar.rieder@i-med.ac.at Web: http://www.icbi.at
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
_______________________________________________ ceph-users mailing list -- ceph-users@ceph.io To unsubscribe send an email to ceph-users-leave@ceph.io
Hi Zakhar, thanks so much for the information. Best D.Rieder On 4/16/24 17:40, Zakhar Kirpichenko wrote:
Hi,
Did you noticed any downsides with your compression settings so far?
None, at least on our systems. Except the part that I haven't found a way to make the settings persist.
Do you have all mons now on compression?
I have 3 out of 5 monitors with compression and 2 without it. The 2 monitors with uncompressed RocksDB have much larger disks which do not suffer from writes as much as the other 3. I keep them uncompressed "just in case", i.e. for the unlikely event if the 3 monitors with compressed RocksDB fail or have any issues specifically because of the compression. I have to say that this hasn't happened yet, and this precaution may be unnecessary.
Did release updates go through without issues?
In our case, container updates overwrite the monitors' configurations and reset RocksDB options, thus each updated monitor runs with no RocksDB compression until it is added back manually. Other than that, I have not encountered any issues related to compression during the updates.
Do you know if this works also with reef (we see massive writes as well there)?
Unfortunately, I can't comment on Reef as we're still using Pacific.
/Z
On Tue, 16 Apr 2024 at 18:08, Dietmar Rieder <dietmar.rieder@i-med.ac.at <mailto:dietmar.rieder@i-med.ac.at>> wrote:
Hi Zakhar, hello List,
I just wanted to follow up on this and ask a few quesitions:
Did you noticed any downsides with your compression settings so far? Do you have all mons now on compression? Did release updates go through without issues? Do you know if this works also with reef (we see massive writes as well there)?
Can you briefly tabulate the commands you used to persistently set the compression options?
Thanks so much,
Dietmar
On 10/18/23 06:14, Zakhar Kirpichenko wrote: > Many thanks for this, Eugen! I very much appreciate yours and Mykola's > efforts and insight! > > Another thing I noticed was a reduction of RocksDB store after the > reduction of the total PG number by 30%, from 590-600 MB: > > 65M 3675511.sst > 65M 3675512.sst > 65M 3675513.sst > 65M 3675514.sst > 65M 3675515.sst > 65M 3675516.sst > 65M 3675517.sst > 65M 3675518.sst > 62M 3675519.sst > > to about half of the original size: > > -rw-r--r-- 1 167 167 7218886 Oct 13 16:16 3056869.log > -rw-r--r-- 1 167 167 67250650 Oct 13 16:15 3056871.sst > -rw-r--r-- 1 167 167 67367527 Oct 13 16:15 3056872.sst > -rw-r--r-- 1 167 167 63268486 Oct 13 16:15 3056873.sst > > Then when I restarted the monitors one by one before adding compression, > RocksDB store reduced even further. I am not sure why and what exactly got > automatically removed from the store: > > -rw-r--r-- 1 167 167 841960 Oct 18 03:31 018779.log > -rw-r--r-- 1 167 167 67290532 Oct 18 03:31 018781.sst > -rw-r--r-- 1 167 167 53287626 Oct 18 03:31 018782.sst > > Then I have enabled LZ4 and LZ4HC compression in our small production > cluster (6 nodes, 96 OSDs) on 3 out of 5 > monitors: compression=kLZ4Compression,bottommost_compression=kLZ4HCCompression. > I specifically went for LZ4 and LZ4HC because of the balance between > compression/decompression speed and impact on CPU usage. The compression > doesn't seem to affect the cluster in any negative way, the 3 monitors with > compression are operating normally. The effect of the compression on > RocksDB store size and disk writes is quite noticeable: > > Compression disabled, 155 MB store.db, ~125 MB RocksDB sst, and ~530 MB > writes over 5 minutes: > > -rw-r--r-- 1 167 167 4227337 Oct 18 03:58 3080868.log > -rw-r--r-- 1 167 167 67253592 Oct 18 03:57 3080870.sst > -rw-r--r-- 1 167 167 57783180 Oct 18 03:57 3080871.sst > > # du -hs > /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/; > iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon > 155M > /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db/ > 2471602 be/4 167 6.05 M 473.24 M 0.00 % 0.16 % ceph-mon -n > mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [rocksdb:low0] > 2471633 be/4 167 188.00 K 40.91 M 0.00 % 0.02 % ceph-mon -n > mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [ms_dispatch] > 2471603 be/4 167 16.00 K 24.16 M 0.00 % 0.01 % ceph-mon -n > mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [rocksdb:high0] > > Compression enabled, 60 MB store.db, ~23 MB RocksDB sst, and ~130 MB of > writes over 5 minutes: > > -rw-r--r-- 1 167 167 5766659 Oct 18 03:56 3723355.log > -rw-r--r-- 1 167 167 22240390 Oct 18 03:56 3723357.sst > > # du -hs > /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/; > iotop -ao -bn 2 -d 300 2>&1 | grep ceph-mon > 60M > /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph03/store.db/ > 2052031 be/4 167 1040.00 K 83.48 M 0.00 % 0.01 % ceph-mon -n > mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [rocksdb:low0] > 2052062 be/4 167 0.00 B 40.79 M 0.00 % 0.01 % ceph-mon -n > mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [ms_dispatch] > 2052032 be/4 167 16.00 K 4.68 M 0.00 % 0.00 % ceph-mon -n > mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [rocksdb:high0] > 2052052 be/4 167 44.00 K 0.00 B 0.00 % 0.00 % ceph-mon -n > mon.ceph03 -f --setuser ceph --setgroup ceph --default-log-to-file=false > --default-log-to-stderr=true --default-log-stderr-prefix=debug > --default-mon-cluster-log-to-file=false > --default-mon-cluster-log-to-stderr=true [msgr-worker-0] > > I haven't noticed a major CPU impact. Unfortunately I didn't specifically > measure CPU time for monitors and , but overall the CPU impact of monitor > store compression on our systems isn't noticeable. This may be different > for larger clusters with larger RocksDB datasets, then perhaps > compression=kLZ4Compression can be enabled by defualt and > bottommost_compression=kLZ4HCCompression can be optional, in theory this > should result in lower but much faster compression. > > I hope this helps. My plan is to keep the monitors with the current > settings, i.e. 3 with compression + 2 without compression, until the next > minor release of Pacific to see whether the monitors with compressed > RocksDB store can be upgraded without issues. > > /Z > > > On Tue, 17 Oct 2023 at 23:45, Eugen Block <eblock@nde.ag <mailto:eblock@nde.ag>> wrote: > >> Hi Zakhar, >> >> I took a closer look into what the MONs really do (again with Mykola's >> help) and why manual compaction is triggered so frequently. With >> debug_paxos=20 I noticed that paxosservice and paxos triggered manual >> compactions. So I played with these values: >> >> paxos_service_trim_max = 1000 (default 500) >> paxos_service_trim_min = 500 (default 250) >> paxos_trim_max = 1000 (default 500) >> paxos_trim_min = 500 (default 250) >> >> This reduced the amount of writes by a factor of 3 or 4, the iotop >> values are fluctuating a bit, of course. As Mykola suggested I created >> a tracker issue [1] to increase the default values since they don't >> seem suitable for a production environment. Although I don't have >> tested that in production yet I'll ask one of our customers to do that >> in their secondary cluster (for rbd mirroring) where they also suffer >> from large mon stores and heavy writes to the mon store. Your findings >> with the compaction were quite helpful as well, we'll test that as well. >> Igor mentioned that the default bluestore_rocksdb config for OSDs will >> enable compression because of positive test results. If we can confirm >> that compression works well for MONs too, compression could be enabled >> by default as well. >> >> Regards, >> Eugen >> >> https://tracker.ceph.com/issues/63229 <https://tracker.ceph.com/issues/63229> >> >> Zitat von Zakhar Kirpichenko <zakhar@gmail.com <mailto:zakhar@gmail.com>>: >> >>> With the help of community members, I managed to enable RocksDB >> compression >>> for a test monitor, and it seems to be working well. >>> >>> Monitor w/o compression writes about 750 MB to disk in 5 minutes: >>> >>> 4854 be/4 167 4.97 M 755.02 M 0.00 % 0.24 % ceph-mon -n >>> mon.ceph04 -f --setuser ceph --setgroup ceph --default-log-to-file=false >>> --default-log-to-stderr=true --default-log-stderr-prefix=debug >>> --default-mon-cluster-log-to-file=false >>> --default-mon-cluster-log-to-stderr=true [rocksdb:low0] >>> >>> Monitor with LZ4 compression writes about 1/4 of that over the same time >>> period: >>> >>> 2034728 be/4 167 172.00 K 199.27 M 0.00 % 0.06 % ceph-mon -n >>> mon.ceph05 -f --setuser ceph --setgroup ceph --default-log-to-file=false >>> --default-log-to-stderr=true --default-log-stderr-prefix=debug >>> --default-mon-cluster-log-to-file=false >>> --default-mon-cluster-log-to-stderr=true [rocksdb:low0] >>> >>> This is caused by the apparent difference in store.db sizes. >>> >>> Mon store.db w/o compression: >>> >>> # ls -al >>> /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph04/store.db >>> total 257196 >>> drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . >>> drwx------ 3 167 167 4096 Aug 31 05:22 .. >>> -rw-r--r-- 1 167 167 1517623 Oct 16 14:00 3073035.log >>> -rw-r--r-- 1 167 167 67285944 Oct 16 14:00 3073037.sst >>> -rw-r--r-- 1 167 167 67402325 Oct 16 14:00 3073038.sst >>> -rw-r--r-- 1 167 167 62364991 Oct 16 14:00 3073039.sst >>> >>> Mon store.db with compression: >>> >>> # ls -al >>> /var/lib/ceph/3f50555a-ae2a-11eb-a2fc-ffde44714d86/mon.ceph05/store.db >>> total 91188 >>> drwxr-xr-x 2 167 167 4096 Oct 16 14:00 . >>> drwx------ 3 167 167 4096 Oct 16 13:35 .. >>> -rw-r--r-- 1 167 167 1760114 Oct 16 14:00 012693.log >>> -rw-r--r-- 1 167 167 52236087 Oct 16 14:00 012695.sst >>> >>> There are no apparent downsides thus far. If everything works well, I >> will >>> try adding compression to other monitors. >>> >>> /Z >>> >>> On Mon, 16 Oct 2023 at 14:57, Zakhar Kirpichenko <zakhar@gmail.com <mailto:zakhar@gmail.com>> >> wrote: >>> >>>> The issue persists, although to a lesser extent. Any comments from the >>>> Ceph team please? >>>> >>>> /Z >>>> >>>> On Fri, 13 Oct 2023 at 20:51, Zakhar Kirpichenko <zakhar@gmail.com <mailto:zakhar@gmail.com>> >> wrote: >>>> >>>>>> Some of it is transferable to RocksDB on mons nonetheless. >>>>> >>>>> Please point me to relevant Ceph documentation, i.e. a description of >> how >>>>> various Ceph monitor and RocksDB tunables affect the operations of >>>>> monitors, I'll gladly look into it. >>>>> >>>>>> Please point me to such recommendations, if they're on docs.ceph.com <http://docs.ceph.com> >> I'll >>>>> get them updated. >>>>> >>>>> This are the recommendations we used when we built our Pacific cluster: >>>>> https://docs.ceph.com/en/pacific/start/hardware-recommendations/ <https://docs.ceph.com/en/pacific/start/hardware-recommendations/> >>>>> >>>>> Our drives are 4x times larger than recommended by this guide. The >> drives >>>>> are rated for < 0.5 DWPD, which is more than sufficient for boot >> drives and >>>>> storage of rarely modified files. It is not documented or suggested >>>>> anywhere that monitor processes write several hundred gigabytes of >> data per >>>>> day, exceeding the amount of data written by OSDs. Which is why I am >> not >>>>> convinced that what we're observing is expected behavior, but it's not >> easy >>>>> to get a definitive answer from the Ceph community. >>>>> >>>>> /Z >>>>> >>>>> On Fri, 13 Oct 2023 at 20:35, Anthony D'Atri <anthony.datri@gmail.com <mailto:anthony.datri@gmail.com>> >>>>> wrote: >>>>> >>>>>> Some of it is transferable to RocksDB on mons nonetheless. >>>>>> >>>>>> but their specs exceed Ceph hardware recommendations by a good margin >>>>>> >>>>>> >>>>>> Please point me to such recommendations, if they're on docs.ceph.com <http://docs.ceph.com> >> I'll >>>>>> get them updated. >>>>>> >>>>>> On Oct 13, 2023, at 13:34, Zakhar Kirpichenko <zakhar@gmail.com <mailto:zakhar@gmail.com>> >> wrote: >>>>>> >>>>>> Thank you, Anthony. As I explained to you earlier, the article you had >>>>>> sent is about RocksDB tuning for Bluestore OSDs, while the issue >>>>>> at hand is >>>>>> not with OSDs but rather monitors and their RocksDB store. Indeed, the >>>>>> drives are not enterprise-grade, but their specs exceed Ceph hardware >>>>>> recommendations by a good margin, they're being used as boot drives >> only >>>>>> and aren't supposed to be written to continuously at high rates - >> which is >>>>>> what unfortunately is happening. I am trying to determine why it is >>>>>> happening and how the issue can be alleviated or resolved, >> unfortunately >>>>>> monitor RocksDB usage and tunables appear to be not documented at all. >>>>>> >>>>>> /Z >>>>>> >>>>>> On Fri, 13 Oct 2023 at 20:11, Anthony D'Atri <anthony.datri@gmail.com <mailto:anthony.datri@gmail.com> >>> >>>>>> wrote: >>>>>> >>>>>>> cf. Mark's article I sent you re RocksDB tuning. I suspect that with >>>>>>> Reef you would experience fewer writes. Universal compaction might >> also >>>>>>> help, but in the end this SSD is a client SKU and really not suited >> for >>>>>>> enterprise use. If you had the 1TB SKU you'd get much longer >>>>>>> life, or you >>>>>>> could change the overprovisioning on the ones you have. >>>>>>> >>>>>>> On Oct 13, 2023, at 12:30, Zakhar Kirpichenko <zakhar@gmail.com <mailto:zakhar@gmail.com>> >> wrote: >>>>>>> >>>>>>> I would very much appreciate it if someone with a better >> understanding >>>>>>> of >>>>>>> monitor internals and use of RocksDB could please chip in. >>>>>>> >>>>>>> >>>>>>> >>>>>> >>> _______________________________________________ >>> ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> >>> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io> >> >> >> _______________________________________________ >> ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> >> To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io> >> > _______________________________________________ > ceph-users mailing list -- ceph-users@ceph.io <mailto:ceph-users@ceph.io> > To unsubscribe send an email to ceph-users-leave@ceph.io <mailto:ceph-users-leave@ceph.io> >
participants (5)
-
Anthony D'Atri
-
Dietmar Rieder
-
Eugen Block
-
Frank Schilder
-
Zakhar Kirpichenko