Dan took care of the boot order last night. Thanks, Dan.
Sadly, the BIOS update was not enough to fix the Dead jobs.
I’ve been experimenting all day with various rebooting techniques to avoid the Dead jobs where a machine fails to reimage. The FOG OS that streams OS images onto the disks just does a
reboot -f.
This can hang for various reasons if there’s a firmware or BIOS bug. Newer hardware makes this a likelihood.
I’ve built a new FOG OS kernel and filesystem and put in place a “magic SysRq reboot” instead (e.g.,
echo b > /proc/sysrq-trigger). Please let me know if you get any Dead jobs after
2026-02-10 19:04:23,097.097.
Last week, I also put in place some logic to automatically change the hostname after a host get reimaged. I was seeing occasional
issues where that would hang, maybe because networking wasn’t fully ready yet. I added some hardening to that and captured new FOG images so I’m
hoping things will be more stable now.
From: David Galloway <David.Galloway@ibm.com>
Date: Monday, February 9, 2026 at 8:38 PM
To: dev <dev@ceph.io>, David Galloway via Sepia <sepia@ceph.io>
Subject: Re: Queue is Paused
Apparently despite using the
—preserve_setting flag with Supermicro’s update utility, the boot order got reset on all of the nodes
so all of the jobs that just got picked up will die. Dan is working on fixing the boot order again. I stopped the dispatcher until this is done.
From: David Galloway <David.Galloway@ibm.com>
Date: Monday, February 9, 2026 at 8:26 PM
To: dev <dev@ceph.io>, David Galloway via Sepia <sepia@ceph.io>
Subject: Re: Queue is Paused
I’ve restarted the dispatcher on soko04. All of the trial testnodes have the latest BIOS and I verified all of the
up and unlocked testnodes reimaged fine with FOG.
Scheduled jobs will continue to be processed now. The hope is this will resolve all of the deployment failures e.g.,
https://tracker.ceph.com/issues/74774 https://tracker.ceph.com/issues/74717.
We are unable to manually reproduce those conditions and we have checked a testnode when it gets left in that condition and we’
re unable to glean any useful information so we are just having to hope
that a BIOS update will take care of it.
From: David Galloway <David.Galloway@ibm.com>
Date: Monday, February 9, 2026 at 12:51 PM
To: dev <dev@ceph.io>, David Galloway via Sepia <sepia@ceph.io>
Subject: Queue is Paused
I’ve paused the teuthology queue so we can update the firmware on the new trial nodes in an attempt to resolve some of the Dead/Fail jobs.
--
David Galloway
Ceph Engineering Labs – Infrastructure Architect
david.galloway@ibm.com
IBM