As you may know, the Sepia Long Running Cluster has been hitting capacity limits over the past week or so. This has resulted in service disruptions to teuthology runs, chacra.ceph.com, docker-mirror.front.sepia.ceph.com, and quay.ceph.io. We've been able to get by by deleting/compressing logs more aggressively but it's not ideal or sustainable. Patrick has created a new erasure coded pool/filesystem that will allow us to keep the same amount of logs but use less space. In order to have teuthology workers start writing logs to that pool, we need to take an outage. At 0400 UTC 19AUG2020, I will instruct all teuthology workers to die after their running jobs finish. At 1300 UTC, I will kill any jobs that are still running. This gives the lab 9 hours to gracefully shut down. At that point, we will switch the mountpoint on teuthology.front over to the new EC pool and start storing new logs there. At the same time, Patrick will start migrating logs on the existing/old pool to the new pool. This means that logs from 7/20 through 8/19 will be unavailable (you'll see 404s) via the Pulpito web UI and qa-proxy URLs until they're migrated to the new EC pool. Let me know if you have any questions/concerns. Thanks, -- David Galloway Systems Administrator, RDU Ceph Engineering IRC: dgalloway
This work is complete. Jobs/workers are writing to the new EC pool. Patrick will begin migrating logs from the last month to the new pool now. Thanks for your patience and thanks Patrick for your help! On 8/17/20 2:24 PM, David Galloway wrote:
As you may know, the Sepia Long Running Cluster has been hitting capacity limits over the past week or so. This has resulted in service disruptions to teuthology runs, chacra.ceph.com, docker-mirror.front.sepia.ceph.com, and quay.ceph.io.
We've been able to get by by deleting/compressing logs more aggressively but it's not ideal or sustainable.
Patrick has created a new erasure coded pool/filesystem that will allow us to keep the same amount of logs but use less space. In order to have teuthology workers start writing logs to that pool, we need to take an outage.
At 0400 UTC 19AUG2020, I will instruct all teuthology workers to die after their running jobs finish. At 1300 UTC, I will kill any jobs that are still running. This gives the lab 9 hours to gracefully shut down.
At that point, we will switch the mountpoint on teuthology.front over to the new EC pool and start storing new logs there.
At the same time, Patrick will start migrating logs on the existing/old pool to the new pool. This means that logs from 7/20 through 8/19 will be unavailable (you'll see 404s) via the Pulpito web UI and qa-proxy URLs until they're migrated to the new EC pool.
Let me know if you have any questions/concerns.
Thanks,
On Wed, Aug 19, 2020 at 8:52 AM David Galloway <dgallowa@redhat.com> wrote:
This work is complete. Jobs/workers are writing to the new EC pool.
Patrick will begin migrating logs from the last month to the new pool now.
The migration is complete. Teuthology artifacts are now taking only 150% space instead of 300%! # ceph df --- RAW STORAGE --- CLASS SIZE AVAIL USED RAW USED %RAW USED hdd 374 TiB 172 TiB 201 TiB 202 TiB 53.92 ssd 2.1 TiB 1.3 TiB 832 GiB 887 GiB 40.50 TOTAL 376 TiB 174 TiB 201 TiB 203 TiB 53.84 --- POOLS --- POOL ID STORED OBJECTS USED %USED MAX AVAIL ... cephfs.teuthology.meta 113 14 GiB 1.08M 41 GiB 4.87 266 GiB cephfs.teuthology.data 114 2.2 TiB 26.00M 6.6 TiB 5.48 38 TiB cephfs.teuthology.data-ec 119 117 TiB 55.12M 184 TiB 61.73 76 TiB Thanks to everyone who helped with this including Josh Durgin and David Galloway. -- Patrick Donnelly, Ph.D. He / Him / His Principal Software Engineer Red Hat Sunnyvale, CA GPG: 19F28A586F808C2402351B93C3301A3E258DD79D
participants (2)
-
David Galloway
-
Patrick Donnelly