Quick update:

Ceph Windows checks on GitHub PRs

This was broken, then fixed yesterday, then broke again overnight.  Some background - I forgot the Windows jobs relied on a qcow2 image that we build and had hosted via filedump.ceph.com from the LRC.  I built this image, deployed a new “apt-mirror” service in the Openshift cluster and restored the lab-extras repo and this Windows qcow2 image to it.

The service originally used an nginx container image from Docker Hub.  Our Openshift cluster enforces image signature/trust policy for external registries, and the pull was rejected, so the nginx pod failed to start and the service went down.  That caused lab-extras to be unavailable resulting in "Failed to download metadata for repo ''lab-extras'': Cannot download repomd.xml: Cannot download repodata/repomd.xml: All mirrors were tried” errors.  The Windows PR checks also failed with "ERROR: The WINDOWS_VM_IP env variable is not set”

This has been resolved by importing and mirroring the nginx image into the internal Openshift registry and pointing the apt-mirror service to use it.  I also modified the Deployment to 3 replicas and added a “Pod Disruption Budget” to prevent all replicas from going down at the same time.

Teuthology Dead jobs

After getting conserver running, I was able to see cloud-init during MaaS provisioning is occasionally timing out trying to pull the testnode’s preseed (kickstart).  This happens under load.  Our current theory is the database is getting hammered and MaaS can’t handle all the concurrent requests so we will be tuning that today.  Hopefully the last few lingering job Dead issues will be cleared up then.

Thanks for your patience.  These last few hurdles are difficult to find until things are under load but we’re nearly there.

From: David Galloway <David.Galloway@ibm.com>
Date: Monday, January 5, 2026 at 4:27 PM
To: dev <dev@ceph.io>, David Galloway via Sepia <sepia@ceph.io>
Subject: Re: Sepia Lab Update

I stand correct on teuthology being ready…  There are still some bugs that need to be worked out in the maas provisioner.  Hold off for now.

From: David Galloway <David.Galloway@ibm.com>
Date: Monday, January 5, 2026 at 2:33 PM
To: dev <dev@ceph.io>, David Galloway via Sepia <sepia@ceph.io>
Subject: Sepia Lab Update

Just wanted to provide a quick update on the lab status.


Overall, you should be able to schedule teuthology runs using the wip-dg-maas teuthology branch from soko04.front.sepia.ceph.com.  You may just need to tolerate some Dead jobs for now as we work through the remaining infra issues.

Apologies for missing the CSC call this morning.  Was not feeling well enough to stare at a screen.

Let me know if there are any questions.

-- 

David Galloway

Ceph Engineering Labs – Infrastructure Architect

david.galloway@ibm.com

IBM