Quick update:
Ceph Windows checks on GitHub PRs
This was broken, then fixed yesterday, then broke again overnight. Some background - I forgot the Windows jobs relied on a qcow2 image that we build and had hosted via filedump.ceph.com from the
LRC. I built this image, deployed a new “apt-mirror” service in the Openshift cluster and restored the lab-extras repo and this Windows qcow2 image to it.
The service originally used an nginx container image from Docker Hub. Our Openshift cluster enforces image signature/trust policy for external registries, and the pull was rejected, so the nginx
pod failed to start and the service went down. That caused lab-extras to be unavailable resulting in "Failed to download metadata for repo ''lab-extras'': Cannot download repomd.xml: Cannot download repodata/repomd.xml: All mirrors were tried” errors. The
Windows PR checks also failed with "ERROR: The WINDOWS_VM_IP env variable is not set”
This has been resolved by importing and mirroring the nginx image into the internal Openshift registry and pointing the apt-mirror service to use it. I also modified the Deployment to 3 replicas
and added a “Pod Disruption Budget” to prevent all replicas from going down at the same time.
Teuthology Dead jobs
After getting conserver running, I was able to see cloud-init during MaaS provisioning is occasionally timing out trying to pull the testnode’s preseed (kickstart). This happens under load. Our current theory is the database is getting hammered and MaaS can’t
handle all the concurrent requests so we will be tuning that today. Hopefully the last few lingering job Dead issues will be cleared up then.
Thanks for your patience. These last few hurdles are difficult to find until things are under load but we’re nearly there.
From: David Galloway <David.Galloway@ibm.com>
Date: Monday, January 5, 2026 at 4:27 PM
To: dev <dev@ceph.io>, David Galloway via Sepia <sepia@ceph.io>
Subject: Re: Sepia Lab Update
I stand correct on teuthology being ready… There are still some bugs that need to be worked out in the maas provisioner. Hold off for now.
From: David Galloway <David.Galloway@ibm.com>
Date: Monday, January 5, 2026 at 2:33 PM
To: dev <dev@ceph.io>, David Galloway via Sepia <sepia@ceph.io>
Subject: Sepia Lab Update
Just wanted to provide a quick update on the lab status.
-
Incoming connections to the lab got resolved as I was typing this e-mail. This means pulpito is back, qa-proxy is back, new VPN connections can be made. Apparently traffic got rerouted to our backup/redundant link and the firewall was not active for that
route. BGP was moved back to the primary link.
-
There are some development machines available for use. They are all installed and assigned at this point so additional users will have to share. I’m trying to keep teammates on the same boxes since you/they are more likely to keep in touch and not step on
each other’s toes. If you need a resource, see
https://wiki.sepia.ceph.com/doku.php?id=hardware:sockeni#users and let me know which machine you need.
-
I got conserver running last week. All of the new testnode (“trial”) machines needed their BMC firmware updated to get serial-over-LAN working. Conserver running means we should have worker_logs being populated during teuthology jobs. However, the dispatcher
doesn’t seem to have picked up that change. We’re looking into it.
-
Speaking of teuthology jobs, there are still a few of the new testnodes giving us trouble. Getting console logs working was a precursor to working out the remaining kinks with deploying the actual hosts.
-
The ‘wip-dg-maas’ branch still has some changes that need to make it back into the
official MaaS PR
-
The Red Hat migrated gear arrived in Poughkeepsie and was unloaded on 12/19. The folio, robsoni, LRC, and Jenkins builders are highest priority to re-rack and bring online at the moment. smithis will be next.
Overall, you should be able to schedule teuthology runs using the wip-dg-maas teuthology branch from soko04.front.sepia.ceph.com. You may just need to tolerate some Dead jobs for now as we work through the remaining infra issues.
Apologies for missing the CSC call this morning. Was not feeling well enough to stare at a screen.
Let me know if there are any questions.
--
David Galloway
Ceph Engineering Labs – Infrastructure Architect
david.galloway@ibm.com
IBM